Synthetic Agent Training Data Works Best as a Closed-Loop System

AutoSynthData turns observed agent failures into validated training tasks. Its larger lesson is that synthetic data quality depends on executable environments, discriminating verifiers, controlled variation, and a curriculum that moves as the model improves.

Source artwork for AutoSynthData: Generating Training Data for Enterprise Agents
Source artwork · Hugging Face / credited contributors ↗
THE SHORT VERSION

Treat synthetic agent data as a closed-loop system of diagnosis, executable task construction, adversarial verification, training, and renewed evaluation.

Synthetic data for agents is easy to describe and hard to make useful. A generator can produce thousands of plausible requests, yet those requests may depend on unavailable tools, conflict with policy, admit no valid solution, or reward an agent that changed the wrong state. The central problem is therefore not volume. It is constructing training examples whose difficulty, environment, solution, and success test agree.

ServiceNow CoreAI’s AutoSynthData release offers a concrete version of that idea. More importantly, it suggests a practical way to reason about agent-data pipelines: treat generation as a feedback-controlled engineering system rather than a one-time content exercise.

What AutoSynthData adds

The released pipeline begins with diagnostic runs in an enterprise environment. It compares failures from a target model with successful behavior from a stronger teacher, converts the resulting gaps into sanitized capability specifications, and generates new tasks rather than reproducing evaluation prompts. Each task combines a system specification, a user request, and a verifier.

Candidates are executed and checked before acceptance. The reported configuration favors examples solved by the target model in no more than one of three trials and by the stronger solver in at least two of three. Positive checks confirm that an intended solution passes; negative checks mutate expected outcomes to test whether incorrect states fail. A bounded critique-and-repair loop can fix candidates, while batch review looks for repetition and missing capability coverage.

In EnterpriseOps Gym experiments, the team reports roughly 2,000 examples per domain. Synthetic supervised fine-tuning improved Hybrid mean Pass@1 by 7.2 percentage points and raised ITSM mean Pass@1 from 18.77% to 27.18%. These are results in the named benchmark environments, not evidence that the same gains transfer automatically to every enterprise workflow.

Think in contracts, not prompts

A useful unit of agent training data is an executable contract. The request says what outcome is wanted. The system specification defines what actions are permitted. The initial state determines what is possible. The verifier defines observable success. If any pair conflicts, the example teaches noise.

Consider a service-desk task: close a duplicate incident, preserve the original incident, and add a cross-reference. A realistic prompt alone is insufficient. The environment must contain both records; the agent needs permission to update the duplicate; and the verifier should check the status, the reference, and the untouched original. Checking only whether one record became closed would reward several wrong trajectories. Conversely, requiring one exact sequence of API calls would reject alternative valid solutions.

That distinction leads to a helpful review question: does the verifier specify the required final properties, or merely imitate the author’s preferred path? State-based checks are often more robust when multiple action sequences can produce the same legitimate result. Trajectory checks remain appropriate for constraints that concern the process itself, such as obtaining approval before a consequential action.

Calibrate difficulty with evidence

Random complexity is not a curriculum. Adding more entities, longer prompts, or arbitrary constraints can make tasks harder without teaching a reusable capability. A better progression varies the dimensions tied to an observed failure: tool selection, ordering, state inspection, exception handling, or respecting a policy boundary.

Repeated trials help distinguish a stable capability gap from sampling noise. If both target and teacher routinely fail, the task may be impossible, ambiguous, or beyond the available tools. If both routinely succeed, it may add little new signal. The productive region is where the target struggles but a reliable solution can still be demonstrated and verified. Trial thresholds should be treated as operational choices, however, not universal constants. They depend on model variance, execution cost, and the consequences of a mislabeled example.

Variation also needs lineage controls. Deriving variants only from vetted core tasks can limit compounding errors. Teams should still compare semantic structure across the batch: changing names and IDs does not create meaningful diversity if every example exercises the same workflow. Coverage should be measured across capabilities and constraints, not only by prompt similarity.

Evaluate the data factory itself

A production pipeline needs metrics beyond accepted-example count. Useful measures include candidate acceptance rate, failure reason, repair success, verifier false-positive probes, capability coverage, duplicate structure, teacher success stability, and cost per accepted example. Sharp changes can expose a broken environment adapter, a permissive verifier, or a generator drifting toward easy cases.

Holdout design matters too. New prompts are not independent if they share templates, states, or workflow skeletons with training tasks. Evaluation should separate task instances and, where practical, reserve capability combinations or workflow families. Human review is especially valuable for realism and policy interpretation, areas where executable checks may be necessary but incomplete.

Finally, the curriculum should move. Once training makes a task family reliably easy, continuing to generate more of it wastes budget and can distort the dataset. Fresh evaluation should identify the next boundary while regression sets preserve earlier gains. The result is a cycle: diagnose, specify, generate, execute, challenge the verifier, train, and diagnose again.

The practical takeaway

AutoSynthData’s most transferable idea is not a particular generator or model pairing. It is the separation of concerns around generation: environment grounding establishes feasibility, teacher execution establishes a workable demonstration, adversarial checks test the verifier, and batch review protects coverage. Each layer catches a different failure mode.

For teams building domain agents, a small collection of executable, carefully challenged tasks can be more informative than a large pile of plausible instructions. Scale becomes valuable only after the contracts are coherent and the feedback loop can show what the current model still needs to learn.

Source: AutoSynthData: Generating Training Data for Enterprise Agents ↗. How we write

← Back to all articles