AutoSynthData turns an enterprise agent’s failures into training data
ServiceNow CoreAI publishes AutoSynthData, a pipeline that converts an enterprise agent’s capability gaps into synthetic training tasks, each described by a system specification, a prompt and a verifier. On the EnterpriseOps Gym benchmark, the fine-tune lifts Pass@1 by 35% relative, without ever touching the original evaluation data.
October 2, 2026. ServiceNow CoreAI publishes AutoSynthData, a pipeline that turns an enterprise agent’s weaknesses into synthetic training data, on the Hugging Face blog. October 2, 2026. The core idea fits in one equation: an agentic task is a triple of “system specification + user prompt + verifier”, and it is the verifier — not the generation — that determines data quality. October 2, 2026. On the EnterpriseOps Gym benchmark, the resulting fine-tune lifts mean Pass@1 by 7.2 points (+35% relative) in the Hybrid domain, and from 18.77% to 27.18% in ITSM. Why it matters: enterprises hit a ceiling that neither more parameters nor more generic data can break through — the key is to manufacture data that targets exactly what their model fails to do.
The agentic task, formalized
The starting point is a formalization. AutoSynthData describes an agentic task as a triple: the system specification (the instructions, policies and initial state, such as a seeded database), the user prompt (what the user asks for), and the verifier (which decides whether the produced trajectory succeeded).
Each component has its own properties. A useful prompt must be feasible (at least one valid trajectory exists in the environment), realistic (it resembles what a user would actually request) and difficult (it exposes a weakness of the model, because an already-solved task provides no signal). The verifier, for its part, must be consistent with the prompt, sound (it rejects trajectories that fail) and complete (it accepts valid solutions without imposing a single one). A lax verifier rewards wrong behavior; an overly restrictive one penalizes correct solutions.
From failures to a curriculum
The pipeline starts from the model’s failures. AutoSynthData evaluates the target model and a stronger teacher in the environment, then identifies where the target fails and how the teacher succeeds. Those observations are distilled into capability specification cards, stripped of anything that could leak: the generator receives neither the original prompts, nor the entities, nor the trajectories, nor the verifier details.
It receives the cards, and uses them to create new tasks — different prompts, different states, different solution paths. In other words, the system does not retrain on the evaluation tests: it learns what to teach, then manufactures unseen exercises that exercise the same capability. As the model improves, the curriculum shifts toward what still resists it.
Generate, then multiply
Dataset construction happens in two phases. The Target phase creates the core set of samples from the capability cards: workers generate tasks in parallel, and each candidate goes through validation, execution, solver evaluation and repair before acceptance. The Multiply phase expands the set by creating novel variants of accepted samples, each with its own request, state, entities and verifier — and each passing the same checks. A variant cannot seed another variant: expansion stays anchored to the vetted core, which limits drift across generations.
On the implementation side, AutoSynthData separates a shared controller — generation, quality control, coverage, dataset construction — from an environment-specific adapter responsible for execution, reference-trajectory replay and deterministic verification.
Verify before you accept
Generating a plausible request is not enough. A task may be impossible in the environment, its reference solution may fail on execution, or its verifier may reward a wrong final state. AutoSynthData therefore checks quality at two levels.
At the sample level, every candidate passes a positive verification (does the intended solution actually solve the task?) and a negative verification (do incorrect outcomes actually fail?). The second gate is crucial: it catches weak verifiers that award success without requiring the intended behavior. Failing candidates go to a critic that diagnoses the cause — inconsistent state, impossible workflow, faulty verifier — and guides a bounded repair rather than restarting generation from scratch.
Difficulty is calibrated with a three-trial rule: a task is only accepted if the target model solves it no more than one time in three, while the stronger solver solves it at least two times in three. That filter keeps samples in the useful zone — hard enough to expose a weakness, solvable enough for the teacher to provide a reliable demonstration.
At the batch level, a meta-review examines the accepted set: which task families are overrepresented, which capability dimensions are missing, which targets keep failing generation. The controller then reduces generation in saturated regions and redirects effort toward the gaps.
The numbers: Hybrid and ITSM
The experiments, run on EnterpriseOps Gym, offer a concrete measure. In the Hybrid domain, with Gemma-4-26B-A4B-it as target and Qwen3.8-27B as teacher, the pipeline generated 2,000 samples in about 18 hours. The best checkpoint (epoch 5) lifts mean Pass@1 by 7.2 points — +35% relative — raises verifier success from 63.01% to 68.55%, and closes 59% of the original gap between Gemma and the reference model.
In the ITSM domain, with DeepSeek-V4.1-Flash as teacher, the pipeline produced 1,994 samples in 66 hours — a slower run that predates the throughput optimizations — and lifted Pass@1 from 18.77% to 27.18%. The method holds in a second domain, ruling out a Hybrid-specific overfit.
The limits and what comes next
AutoSynthData remains a targeted demonstration, and its authors set the boundaries themselves. The experiments cover SFT (supervised fine-tuning); the same mechanism could feed reinforcement learning — generate tasks that challenge the current policy, train, then move the generation target with the updated policy — but that is still an announced direction rather than a result.
The deep difficulty is not text generation but verification. Writing a sound and complete verifier for enterprise tasks — where “success” depends on a database state, a business policy or a tool chain — is an open problem. AutoSynthData works around it with negative verification and assisted repair, but the final quality of the fine-tune will remain capped by the quality of the verifiers an enterprise can write.
The release fits a broader shift: after the era of “large-scale” synthetic data of the self-instruct kind, the field is moving toward executable and verifiable data, grounded in a real environment. AutoSynthData is the enterprise flavor of that idea — data generated not to look like text, but to be run and judged.
The practical takeaway is that synthetic data is no longer a quantity game. 2,000 well-verified samples — a fraction of the datasets used in generic fine-tunes — moved the needle by 7.2 points because each one was executable, realistic and verifiable. For enterprise teams, that reframes the effort: the budget is better spent on writing a few hundred solid verifiers than on generating millions of unverifiable prompts.
Verdict
If your enterprise agents plateau despite bigger models, the lever is not yet another fine-tune on generic data: it is a failure-driven pipeline that manufactures targeted synthetic tasks, each validated by a sound and complete verifier. The real bottleneck is not generation but the verifier: it decides whether your training data teaches the right behavior or rewards shortcuts. If you are starting from scratch, begin by formalizing your tasks as a “specification + prompt + verifier” triple and invest in the positive and negative gates before multiplying volume — AutoSynthData’s 35% comes from the quality of the checks, not the quantity of text.