A candidate model improves average answer quality. The team asks whether to release it. The average does not say whether it now leaks restricted information, needs more review, or breaks a task that a small user group depends on.
A release gate should describe a decision under evidence and constraints. It is not a ceremony where a single score turns green. The useful artifact lets another person understand what changed, what was tested, what remains uncertain, and who can stop the release.
Start with the supported promise
For a synthetic policy assistant, define the user group, source universe, allowed actions, required version checks, and fallback. The release contract might permit source-backed explanations but prohibit changing employee records. A candidate that drafts better prose does not earn new write authority.
Keep scope visible. Results on one language or document class do not establish support for another. A new model version can improve the existing task without justifying a broader product promise.
Separate blockers from trade-offs
Some properties can be evaluated as rates with tolerable limits; others may be release blockers within the tested scope. For example, any observed cross-account data exposure deserves investigation, regardless of a better average relevance score.
A zero count on a finite suite is not proof that the event cannot happen. Record test coverage and the assumptions that make the operating scope acceptable. Thresholds should be chosen before seeing the candidate's score, with the accountable product and domain owners.
Preserve the comparison
Run the current and candidate configuration on the same versioned cases and starting states. Keep prompt, model, tools, retrieval snapshot, validation rules, and judge rubric with the report. Inspect changed examples and record repeated trials where stochastic variation matters.
Official evaluation guidance emphasizes task-specific criteria and calibration against human judgment. That supports a layered test: code for deterministic invariants, source-aware judgment for semantic criteria, and actual task completion at the user boundary.
Write a compact decision table
| Question | Evidence to retain |
|---|---|
| What changed? | Effective configuration and intended behavior |
| Did supported tasks improve? | Paired outcomes, denominators, important slices, uncertainty |
| What could block release? | Critical failures and tested constraints |
| How will it roll out? | Exposure, monitoring window, owner, stop conditions |
| How will it recover? | Configuration rollback, state reconciliation, user communication |
The table is a proposed operating artifact, not a validated universal standard. Keep it small enough that reviewers can inspect the underlying cases instead of only approving labels.
Online exposure answers another question
Offline evaluation cannot reproduce every production source, user, and dependency. A canary can reveal operating behavior under limited exposure, but it needs a traffic assignment, observation window, and stopping authority. Low exposure limits reach; it does not make an unacceptable action permissible.
The SRE canary guidance describes controlled release comparisons. For agents, isolate or disable writes during shadow evaluation so the candidate does not duplicate real effects while being “invisible.”
Include the state after release
Restoring a previous configuration does not undo an email already sent or a record changed. Describe which effects can be reversed, which need compensation, and which require notifying affected users. The release owner needs access to those recovery paths, not only a model-version dropdown.
The NIST GenAI Profile can organize risk and responsibility, but it does not certify your specific gate. The final decision should state the evidence-supported scope and the conditions that would require revisiting it.