Robot Benchmarks Need to Diagnose Failures, Not Just Rank Policies

ProjectSim argues that robotics evaluation should connect observed failures to targeted training data. The practical challenge is designing benchmarks that separate perception, control, recovery, and real-world robustness while keeping simulation honest.

Source artwork for Robotics benchmarking: where we are, and what comes next
Source artwork · Hugging Face / credited contributors ↗
THE SHORT VERSION

Robotics benchmarks become more actionable when they expose failure stages, specify the type of novelty, test recovery in closed loop, and verify that simulated practice transfers to hardware.

Robotics has no shortage of leaderboards, but a higher score rarely tells a team what to collect, simulate, or change next. A new ProjectSim overview makes that gap the central issue. It surveys benchmarks ranging from component tests to long-horizon household tasks, then proposes a tighter loop: evaluate a policy, identify a repeatable failure, add experience aimed at that failure, and test whether the intervention improves real behavior.

That proposal matters because robot learning spends expensive resources whenever it turns a diagnosis into new data. A benchmark should therefore do more than sort policies. It should make the next experiment easier to choose and harder to misinterpret.

One score can hide several different systems

A manipulation result is produced by a chain of capabilities. The robot must perceive the scene, interpret the request, choose an action, handle contact, notice unexpected motion, and recover. A binary task-success metric compresses that chain into a useful final outcome, but it cannot locate the weak link.

The ProjectSim article illustrates the range of questions already covered by existing benchmarks. Some isolate pose estimation or grasp selection. Others execute full policies in simulation, test sequences of language-directed tasks, vary camera or scene conditions, or use physical hardware. It also highlights results in which policies performed individual skills more reliably than combinations of skills, and cases where apparent instruction following survived even when the instruction was removed. Those examples show why success alone can overstate the ability a test intends to measure.

A better report keeps the end-to-end score while adding structured intermediate outcomes. For a pick-and-place task, evaluators might record whether the correct object was selected, whether a stable grasp occurred, whether transport succeeded, and whether placement met the tolerance. The stages should be observable and task-relevant, rather than internal signals chosen because they are convenient to log. This produces a failure profile without redefining success downward.

Evaluation splits must name the kind of novelty

The word “generalization” is too broad for a robot benchmark result. An unfamiliar mug in a familiar kitchen is a different test from a familiar mug viewed by a moved camera. Both differ again from a novel instruction that combines known skills. If these cases share one aggregate, users cannot tell what kind of deployment change the policy can tolerate.

Useful evaluation matrices vary one factor at a time before testing combinations. A minimal design can include object identity, object pose, background and lighting, camera geometry, instruction wording, and task composition. Training exclusions should be documented at the same level. Otherwise an “unseen” test may differ only cosmetically, or it may quietly require an entirely new interaction.

This matrix also guards against shortcut learning. If changing the instruction while holding the scene fixed does not change the policy’s target, the system may be following scene regularities rather than language. If a camera shift causes collapse while object poses remain stable, the weakness is more likely visual than mechanical. Controlled contrasts turn a surprising failure into a hypothesis that another experiment can challenge.

Recovery deserves its own protocol

Physical tasks do not unfold from pristine starting states. Objects slip, parts catch on edges, and earlier actions leave the workspace altered. A policy that succeeds from a standard reset can still fail as soon as it creates its own off-nominal state.

Recovery evaluation should include at least two modes. In the first, the policy begins directly in a curated failure state. This tests whether it recognizes and escapes a known problem. In the second, it runs from the normal initial state and must deal with its own errors. This measures whether recovery skills appear at the right moment in the full closed loop. ProjectSim cites research where recovery examples helped in injected failure states but offered little improvement on complete attempts, demonstrating why the distinction is consequential.

The training comparison must also be fair. If a team adds recovery demonstrations, it should compare them with an equal data or interaction budget spent on ordinary task practice. Report recovery-stage success and full-task success separately. An intervention that fixes insertion after misalignment but damages initial grasping is not an unqualified gain.

Simulation needs behavioral fidelity, not visual resemblance alone

Simulation is attractive because trials can be reset and varied cheaply. Yet a visually convincing reconstruction may erase the very interaction that needs practice. Geometry near contact points, friction, mass, compliance, latency, and sensor noise can change whether a grasp holds or an insertion jams. Background appearance can matter too when observations feed a vision-based policy, but matching pixels is not sufficient.

ProjectSim’s stated experiment reconstructs physical scenes, assigns physical properties, runs the same policies in real and simulated tasks, and compares both final scores and the stages reached. The important test is agreement in behavior: do policies fail in similar places, and does the simulator preserve their relative strengths and weaknesses? A simulator that raises every score may be easy rather than faithful. A simulator that matches an average while producing different failure modes may teach the wrong recovery.

Teams can evaluate fidelity with paired evidence rather than a single realism label. Compare success rates with uncertainty, stage-transition rates, failure categories, and sensitivity to controlled changes. Validate across more than one policy, because tuning a scene until one controller matches hardware can overfit the reconstruction to that controller. Hold out some tasks or policies when adjusting simulation parameters, then use them as a stronger test of transfer.

A benchmark should end with a decision

The most valuable benchmark output is a bounded next action. That might be collecting examples of withdrawing and realigning a stuck part, adding camera variation, improving task annotations, or rejecting a simulator until contact behavior is closer to hardware. Each action should correspond to an observed failure and a follow-up metric.

This turns evaluation into an experimental cycle rather than a ceremony at the end of training. The discipline is simple: preserve end-to-end success as the ultimate measure, expose enough stages to diagnose failure, declare exactly what is held out, compare targeted data with a budget-matched baseline, and confirm improvements on real hardware. Rankings still matter, but diagnosis is what helps the next robot become better.

Source: Robotics benchmarking: where we are, and what comes next ↗. How we write

← Back to all articles