TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
ServiceNow CoreAI’s AutoSynthData pipeline uses a target model’s failures and a stronger teacher’s successes to generate new enterprise-agent training tasks. The Hugging Face report describes how the system grounds and validates tasks in an environment, but does not provide performance results or establish that the method works across enterprise settings.
ServiceNow CoreAI has described AutoSynthData, a system designed to turn enterprise-agent failures into new training tasks that can be checked in the environment where the agents operate. In a report published on Hugging Face, the team says the pipeline uses a target model’s evaluation results and a stronger teacher model’s successful runs to identify capability gaps, generate varied tasks and validate them before using accepted samples for post-training.
The report frames an agent task as three linked parts: a system specification defining instructions and environment constraints, a user prompt describing the requested work, and a verifier that checks whether the final result meets the task requirements. The authors say useful generated tasks must be feasible with the available tools and state, realistic for the environment, and difficult enough to expose a weakness in the target model.
AutoSynthData begins by evaluating the target model on diagnostic tasks. The report says a stronger teacher model is also run on the evaluations to help characterize which tasks can be solved and what successful behavior involves. The team examines capabilities, tool use, workflow structure, failure points, correct end states and the ways a task can vary without changing the skill being tested.
Those observations are distilled into what the report calls sanitized capability specification cards. According to the authors, the task generator receives these cards rather than the original evaluation prompts, entities, trajectories or verifier details. It uses the cards to produce new tasks, checks them in the environment and sends accepted samples into post-training. Afterward, evaluation of the updated model can inform another generation round. The report illustrates the pipeline using EnterpriseOps Gym and a released dataset.
Training Agents on Enterprise Workflows
Enterprise agents must act within the particular systems, policies, tools and data states of each organization. A model can perform well on broad tasks yet mishandle a particular workflow or fail to respect a local constraint. AutoSynthData addresses the gap between identifying such a weakness and assembling enough varied examples to train against it.
The approach also makes task checking part of the training-data process. A verifier that accepts incorrect work could reward unsafe or incomplete behavior; one that is too restrictive could reject valid solutions. The report therefore treats verifier quality as a central design requirement, alongside task generation. If the pipeline can produce realistic tasks with reliable checks, it could help teams target training toward operational weaknesses rather than relying only on general-purpose examples. The report describes this potential, but does not establish measured gains in real deployments.
enterprise AI training data generation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Evaluation to New Tasks
The method starts from an existing environment: the world an agent can observe and change, the tools and APIs available to it, and the effects of its actions. In that setting, an individual failure can reveal a problem, but the report argues that training requires additional tasks that test the same capability under different circumstances.
AutoSynthData’s proposed cycle connects evaluation, task generation and post-training. The target model’s evaluation failures identify candidate gaps; a stronger teacher’s behavior helps describe successful performance; capability cards guide the creation of new tasks; and environment-based checks filter the resulting samples. The team says the original diagnostic tasks are not simply copied into generation, which is intended to produce different prompts, states and solution paths while retaining the relevant capability.
The report presents the process through EnterpriseOps Gym, citing work by Malay and colleagues dated 2026 and pointing to a released dataset. The supplied report excerpt does not include quantitative results, a comparison with other data-generation methods or evidence from production environments.
“At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data.”
— ServiceNow CoreAI, in the Hugging Face report
AI model evaluation and validation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evidence and Deployment Still Open
The supplied material does not report how much AutoSynthData improves an agent, how its results compare with alternative training-data approaches, or how many generated tasks were accepted. It also does not specify whether the system is deployed in production, which enterprise settings beyond the EnterpriseOps Gym demonstration have been tested, or how much human review is required.
Further details about the generation process are incomplete in the provided excerpt, including the full workflow example and the scale and diversity of the generated dataset. The report’s claims describe the authors’ design and experiment; they should not be treated as independent evidence of performance across enterprise environments.
automated machine learning training datasets
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Further Results to Watch
The report says evaluation of an updated model can reveal which capability gaps remain and guide another round of task generation. That creates a proposed iterative process, but the supplied source does not state a schedule for further experiments or announce a product release.
Readers assessing the approach will need additional results showing task acceptance rates, verifier reliability, changes in target-model performance and comparisons against other training methods. Evidence across different enterprise environments would also help establish whether the pipeline generalizes beyond the EnterpriseOps Gym example.
enterprise workflow automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is AutoSynthData?
AutoSynthData is a system described by ServiceNow CoreAI for generating and checking enterprise-agent training tasks based on capability gaps found in model evaluations.
How does it identify what an agent should learn?
The report says it evaluates a target model on diagnostic tasks, compares its runs with a stronger teacher’s behavior, and distills the findings into capability specification cards that guide generation.
How does the system check generated tasks?
Each task includes a verifier intended to determine whether the agent’s outcome satisfies the prompt and system constraints. The authors say tasks are checked in the environment before accepted samples are used for post-training.
Has AutoSynthData been shown to improve production agents?
The supplied report material does not provide production results or quantitative performance comparisons. It describes an experiment using EnterpriseOps Gym and a released dataset.
What remains unknown?
Reported performance gains, verifier accuracy, task-generation scale, human-review needs and results across other enterprise environments are not specified in the material provided.
Source: rss
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
