An agent fixes one support ticket beautifully. On the next attempt, with the same starting state, it selects a different account. A third attempt reaches the right result after an unnecessary tool call. “It can do the task” and “it reliably does the task” are different statements.
That distinction is receiving explicit research attention. Rabanser and colleagues' 2026 reliability study separates consistency, robustness, predictability, and safety. Its evaluations use two benchmarks and particular scaffolds; the results are a reason to inspect these dimensions, not a universal forecast of every agent's production behavior.
Two repeated-run questions
pass@k asks whether at least one of k attempts succeeds. It is useful when a process can afford several candidates and has a trustworthy way to select a correct one. pass^k, introduced for repeated reliability in τ-bench, asks whether all k attempts succeed. Neither should be reported simply as “accuracy.”
For a synthetic task with independent attempts and constant success probability p=0.8, at least one success in three attempts has probability 1−(1−p)^3=0.992. All three succeeding has probability p^3=0.512. Those are consequences of the assumptions, not observed benchmark results. Actual tasks have different difficulty and failures may be correlated. Estimate the metric from repeated trials rather than applying one average p to an entire workload.
Reset the world between trials
A booking task needs the same calendar, user identity, tool permissions, and available slots at the start of each independent trial. If the first attempt books the slot, the second attempt is solving a different task unless the environment is reset. Keep the task fixture and model configuration versioned.
Define success in the environment: the intended booking exists, no conflicting booking was added, and approval was respected. A final message claiming success is not an outcome validator. Conversely, several valid action sequences can reach the same goal; do not penalize harmless sequence variation unless order is part of the contract.
Test more than consistency
Add controlled changes: an irrelevant document, a tool timeout, a missing optional field, or an alternative wording of the same request. These probe robustness. Test whether the system recognizes missing evidence or declines an unsupported action; that probes whether its limits are predictable. Record the severity of failures separately. A missed suggestion and an unauthorized write should not disappear into the same compensating average.
There is no finite perturbation suite that proves general robustness. State what was varied, what stayed fixed, and which failure families remain untested. Model-generated confidence is not a calibrated estimate of task success.
Keep the operating cost in view
A best-of-three process pays for unsuccessful candidates too. It also needs verification, which can be expensive or wrong. Record total tokens, elapsed time, tool calls, retries, and human review across the complete request. For a state-changing task, retrying every candidate is not an acceptable search procedure unless each run is isolated.
A useful release table therefore has separate columns for first-run success, repeated success, constraint violations, recovery outcomes, and cost per accepted task. Report task counts and repeat counts. Preserve pairing when comparing releases, and use a task-level uncertainty analysis that respects repeated observations.
A small first reliability suite
Start with a bounded workflow and a frozen set of ordinary, ambiguous, and failure cases. Repeat each from the same initial state. Inspect every consequential failure, then ask whether the fallback completed the user's task. Expand coverage based on observed gaps, while keeping fresh release cases separate from development cases.
The goal is a precise operating claim: under these conditions the agent completed these tasks this consistently, with these remaining failures. That is more useful than a demonstration that happened to work once.