A model judge rates a candidate answer more highly than the current answer. The score is tempting to treat as evidence of product improvement. But the judge may prefer the longer answer, the first answer, or an answer that repeats its own phrasing while missing a factual error.
A judge is another model-based component. Its outputs need validation against the criterion you actually care about, not only agreement with another automated scorer.
Define the judgment narrowly
For a synthetic research summary, separate claim support, preserved qualifications, source coverage, and readability. A single “quality out of ten” invites different interpretations. Ask the judge to identify the relevant source passage and explain a specific failure category.
Use deterministic checks for properties code can establish: required fields, known citation identifiers, allowed actions, and final state. A language model should not be the only validator of a database write or access decision.
Known biases remain test questions
Zheng and colleagues' LLM-as-a-judge work examines position, verbosity, and self-enhancement biases in the studied settings. It also reports substantial agreement with human preferences. That agreement concerns those models, data, and criteria; it is not a universal threshold for judging enterprise factual correctness.
Reverse answer order, compare concise and verbose versions with equivalent content, and test incorrect answers written confidently. A judge that changes preference when order changes has a measurement problem worth reporting.
Give the judge the necessary evidence
A judge cannot verify a current policy from an answer alone. Supply the authorized source and applicable version, or define the result as a preference judgment rather than a factual check. Missing reference material should produce an explicit insufficient-evidence state.
Do not put untrusted candidate text into a judge instruction field. An answer can contain instructions aimed at the grader. Treat it as data and test injection attempts, especially when judge scores trigger downstream actions.
Build an independently reviewed set
Have qualified reviewers label representative correct, incorrect, ambiguous, and incomplete cases under the same rubric. Include serious errors that are rare in normal traffic. Preserve disagreement and use adjudication where consequences warrant it.
Agreement is one measure. Also examine false acceptance of critical errors, unnecessary rejection of good answers, and performance by failure type. Human labels are not infallible; the review protocol and source access affect their quality too.
Version the measuring instrument
Keep judge model, prompt, rubric, reference set, sampling settings, and aggregation rule with the evaluation artifact. A score change after modifying the rubric may be a measurement change rather than an application improvement.
Official OpenAI evaluation guidance recommends task-specific evaluation and human calibration of automated scoring. Apply that principle without assuming that any particular named model remains the best judge for your workload.
Use scores to inform, not hide, a decision
Compare candidates on paired tasks, inspect changed examples, and repeat cases where judge variability could alter the conclusion. State the sample size and uncertainty method. Treat critical access or effect violations as separate blockers rather than allowing a high writing score to compensate for them.
A useful judge saves review effort while preserving the ability to find mistakes. Its authority should be proportional to what the evaluation has actually demonstrated.