AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Agents Are Learning Inside Your Software. Check Ironclad’s Fine Print on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training GPT-6 Astra in hosted copies of Ironclad’s contract-management software using 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria, while estimated completion times were simulated; the results do not establish deployment-ready performance or measured customer savings.

OpenAI said it trained its GPT-6 Astra model on tasks performed in hosted copies of Ironclad’s contract-management software, reporting that Astra met an average 55% of evaluation criteria across 11 legal, commercial and procurement tasks. The results describe a research test, not verified customer productivity gains: OpenAI says its time estimates were simulated, and its post says human oversight remains necessary.

The tasks were selected by Ironclad staff and OpenAI employees who use the product. They included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes for each task. Evaluation used rubrics with 8 to 50 criteria, depending on the task’s complexity.

OpenAI reported that GPT-5.6 Sol, run at a high setting, met an average 41.6% of criteria, while GPT-6 Astra, run at a maximum setting, met 55%. An internal model used in Astra’s development reached 63.7%. Astra met about 94% of the criteria on one showcased task. These are shares of rubric criteria, not percentages of tasks fully completed or passed.

For the training data, OpenAI said it created synthetic tasks from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information filtered out. The company said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data. The models practised in hosted copies of Ironclad’s product, according to the post.

At a glance
reportWhen: Published October 6; the source does no…
The developmentOpenAI published results from training and evaluating a frontier model on workflows inside Ironclad’s contract-management software and invited more software companies to propose similar research partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Scores Matter

The reported score points to both progress and a practical limit. Contract and procurement workflows often depend on several rules being followed together: for example, a purchase may require Finance approval above a threshold, Security review for certain requests and Legal review when terms are nonstandard. Missing one required step can undermine the workflow even if the agent handles the rest correctly. A 55% average across evaluation criteria does not show which requirements failed on every task, nor does it establish that the agent can safely run those processes without review.

OpenAI’s simulated timing figures also should not be read as demonstrated savings. It estimated 19.2 minutes per attempt for Astra, compared with 37 minutes for GPT-5.6 Sol, but said those figures rely on assumed processing and generation speeds. They are not measured customer times, and the test covered only the 11 research tasks. The published comparison does not show how long a person would take to check and correct an agent’s work.

For businesses buying or operating software, the partnership model may affect more than automation. Training agents in a product could make them better at operating its workflows, while increasing the importance of the software’s underlying business rules, permissions, records and audit controls. That is a potential implication, not a demonstrated market outcome. The test offers no evidence yet about how customers will use such agents in live systems or whether vendors will retain control of the customer experience.

Amazon

contract management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Product Tests to Agent Training

The October 6 post, titled “Advancing computer use with Ironclad,” describes a collaboration with a contract-management software company. The title led some AI news trackers to treat Ironclad as an agent framework, but the source identifies it as the software vendor whose product hosted the test workflows.

OpenAI framed the work as an effort to train models to understand business rules, complete multi-step tasks in specialized software and check whether the result meets the original requirements. The post also invites a small number of software companies to propose research partnerships. It asks potential partners to bring a concrete example of a task current agents cannot reliably complete, people with detailed knowledge of the work, a secure test environment and data that can be used safely for research.

The report is a limited evaluation, not a general benchmark of contract software or a claim that the model can perform all legal work. Its findings apply to the selected tasks and evaluation method described by OpenAI. The post’s own account acknowledges that losing track of a business rule can constrain what a company can confidently assign to an agent.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Test Does Not Establish

The post does not provide enough detail to determine how Astra performed on each rubric requirement across all 11 tasks, how often it produced a result that would be acceptable without correction, or how much time human review would add. An average 55% criteria score can conceal important differences between tasks and between individual requirements.

It is also unclear whether the results will translate to live customer environments, how performance changes with different contracts or company rules, or what safeguards would govern any future deployment. OpenAI says it used no non-public Ironclad customer data in this research, but the source does not describe the full security arrangements for future partnerships, data-retention terms, or the division of responsibility if an agent makes an error.

The time comparison remains an estimate rather than an observed productivity result. The source material also does not specify the year of the October 6 publication, and it gives no schedule for additional partnerships or follow-up results.

Amazon

AI-powered procurement software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Questions Before Agents Go Live

OpenAI says it is seeking a small number of software-company partners, but it has not announced a timetable or named further participants in the material provided. Any next research results will need to clarify the tasks tested, the criteria missed and how performance is measured before readers can judge whether the approach is suitable for consequential business workflows.

Companies considering agents in contract, finance or customer-record systems can ask vendors for the full evaluation criteria and task-level failures, rather than relying on a single average score. They can also ask what approval steps remain mandatory, who reviews outputs, how actions are recorded and what data is used to train or evaluate the system. For now, the reported findings support further testing; they do not establish that agents can replace human review in contract workflows.

Amazon

AI contract analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI test with Ironclad?

OpenAI tested models on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. The tasks included setting up nondisclosure agreements and building procurement approval processes.

What does Astra’s 55% score mean?

It is the average share of evaluation criteria met across the tasks. It is not the percentage of tasks completed successfully, and it does not mean that the resulting workflows are safe to use without checking.

Did the test prove that customers will save time?

No. OpenAI said its time figures, including an estimated 19.2 minutes per Astra attempt, were simulated using assumed processing and generation speeds. They were not measured customer savings.

What data did OpenAI say it used?

OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Can businesses use these agents without human review?

The results do not support that conclusion. OpenAI’s post says human oversight still matters, and the average score indicates that the model missed a substantial share of the evaluation criteria in the reported test.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Memento Constraint: Why Continual Learning Is the Trillion-Dollar Bottleneck Nobody Is Pricing

Exploring how the inability of current AI models to learn continually could reshape the trillion-dollar enterprise AI economy, with insights from recent research.

The New Face Of AI: Anthropic’s Claude Fable 5.1 And Mythos 5.1 Explained

Anthropic announces two new products, Claude Fable 5.1 and Mythos 5.1, but details on capabilities, availability, and use cases remain unclear.

Corvus ISR Day 1: Crafting A WAMI Exploitation Stack Using Synthetic Data

Corvus ISR launches with a synthetic WAMI scene featuring live detection and tracking, demonstrating a new approach to exploitation software for wide-area motion imagery.

AI Grammar and Style Editors to Enhance Quality

Keen on perfecting your writing? Discover how AI grammar and style editors can transform your work and elevate your communication.