Reliable Agents Need State Checks, Not Confident Sign-Offs
ThinkingBox evaluates whether an AI agent leaves business systems in the required final state, then repeats each workflow to expose inconsistency. Here is how that changes agent testing and deployment decisions.

Judge tool-using agents by verified final state and repeated success, then price the full operating policy—including checks, retries, and escalation—not the model call alone.
A polished final answer is weak evidence that an AI agent completed a job. The durable evidence lives elsewhere: the ticket status, the refund record, the booking, the audit log, and every unintended side effect. That distinction matters whenever a model can use tools to change a real system.
Microsoft and Hugging Face introduced ThinkingBox to evaluate this gap. Its benchmark runs 507 synthetic, stateful business workflows across retail, insurance, travel, banking, and consulting. Each task starts from isolated backend state, gives an agent an MCP tool surface, and judges the state left behind. Each model-task pairing is repeated 20 times, making consistency visible rather than treating one successful run as sufficient proof.
Outcome checks change what counts as success
Many agent evaluations focus on the transcript. They ask whether the response sounds helpful, whether a tool call was syntactically valid, or whether the agent selected a plausible function. Those checks are useful, but they observe the path rather than the destination.
A state-based evaluator asks stricter questions. Was the intended record changed? Does it contain the required value? Were all required side effects produced? Did the agent modify anything it should have left alone? A run can therefore fail even if every tool call returned without an error and the closing message sounds certain.
This suggests a layered evaluation design. Use schema validation to catch malformed calls, trajectory inspection to diagnose behavior, and terminal-state assertions to decide whether the task actually passed. Natural-language grading still has a role for requirements that cannot be represented cleanly in a database, such as whether a mandatory disclosure appeared. But whenever a machine-verifiable invariant exists, it should carry more weight than the agent's own account of its work.
Repetition separates capability from dependability
A single-attempt success rate answers how often an agent succeeds across sampled runs. It does not answer whether a particular workflow is safe to hand over. ThinkingBox therefore distinguishes tasks solved at least once from tasks passed in all 20 recorded attempts. The gap can be large: the announcement reports that Kimi-K3 solved 476 of 507 tasks at least once, while only 68 passed in every attempt. Claude Opus 5 passed fewer tasks at least once but completed 241 tasks in all 20 trials.
Neither view is universally sufficient. Best-of-many coverage can matter during exploration, where a human selects or repairs an output. Every-run consistency is more relevant when an agent can commit a payment, close a support case, or update a regulated record without review. Teams should name the metric they optimize instead of compressing these different operating goals into one leaderboard number.
Repeated evaluation also reveals whether failures cluster around specific tasks. An aggregate score can hide a workflow that works nine times out of ten but fails unpredictably on the tenth. For a high-impact action, that remaining uncertainty may be the principal design constraint. The practical response could be stronger verification, a narrower tool set, an approval step, or refusing automation for that workflow.
Build the verifier beside the agent
A useful state check begins with explicit postconditions. For a ticket workflow, those might include the correct status, an attached note, an unchanged customer profile, and no duplicate ticket. The evaluator should reset the environment before each attempt and inspect authoritative storage afterward. This avoids rewarding a model for merely saying it performed an action.
The same pattern can be applied in production:
- Define required and forbidden changes before writing the prompt.
- Give the agent only the tools needed for that workflow.
- Read the relevant records again after the agent finishes.
- Compare observed state with the postconditions.
- Commit, retry, escalate, or compensate according to a typed failure policy.
Retries deserve particular care. Repeating the full workflow after a timeout can create duplicate side effects. A safer design makes operations idempotent where possible, assigns stable operation identifiers, and distinguishes a tool error from an unknown result. If the system cannot establish whether a mutation happened, it should reconcile state before attempting it again.
Verification also needs independence. If the same model both acts and informally declares its own result correct, shared blind spots remain. Deterministic database assertions are preferable for exact facts. A separate narrow rubric can handle semantic requirements, with its limitations documented. Logs should preserve the proposed action, tool response, final state, and verdict so failures can be reproduced rather than debated from screenshots.
Price the reliability target, not only the model call
ThinkingBox frames cost in two ways: cost per successful attempt and cost per task that passes all 20 runs. That distinction is useful beyond this benchmark. A cheap model can become expensive when failures trigger retries, manual review, reversals, or customer support. Conversely, the most capable model may be unnecessary when a constrained workflow has strong deterministic checks and low-cost recovery.
A deployment estimate should therefore include inference, verification, retry probability, human escalation, and the cost of an incorrect irreversible action. Evaluate complete operating policies—model plus tools plus safeguards—rather than choosing a model from pass rate alone.
The larger lesson is straightforward: an agent's response is a claim about completed work. Reliable automation requires independent evidence in the system of record, measured repeatedly under clean conditions. ThinkingBox makes that evidence the center of the evaluation, which is where production agent testing should begin.
Source: The Agent Said It Was Done. The Database Disagreed. ↗. How we write


