AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Put AI Agents Through Tough Scenarios Before Business Use on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate, a live experiment from ThorstenMeyerAI.com, ran five frontier AI models through a simulated software company’s worst week. All models detected every crisis and refused manipulation attempts, but only two closed a €55,000 deal their own analysis justified — and the enterprise pilot extends the wargame to companies’ own read-only data.

The final Crucible League, completed in July 2026, put five frontier AI models in charge of the same small software company during its worst week — and found that spotting every crisis and refusing every manipulation attempt was not enough to run a business well. Only two of the five models signed the €55,000 deal their own analysis had earned, according to results published by Firmulate, the live AI-agent wargame experiment run by ThorstenMeyerAI.com. The project is now offering enterprise pilots that run the same wargame against a read-only export of a company’s own data.

In the final standings, gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Every decision the models made was versioned and auditable, and partial progress counted toward scores. However, the scoring rule capped results on a single principle, in the experiment’s own words: “no amount of good work outweighs a breach of trust.”

The headline result was not missed emergencies. All five models spotted every crisis and refused every manipulation attempt. The divide came after diagnosis: three of the five failed to close a deal their analysis supported. As the experiment puts it, “Same diagnosis, same pitch — no signature.” The decisive competitive weakness was buried two document references deep in the company’s own files, not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

Trust was tested separately. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Thoroughness did not guarantee a strong finish: Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last after leaving the close on the table and attempting to write into a locked department instead of escalating. A weaker version of that boundary weakness appeared in all four other models. Firmulate also notes a fairness caveat: Kimi K3 ran at the API default effort setting while the others ran at xhigh.

At a glance
reportWhen: Crucible League completed July 2026; en…
The developmentFirmulate has published final results of its July 2026 Crucible League and is offering enterprise pilots that run the same wargame against a read-only export of a company’s own data.
Put AI Agents Through Tough Scenarios Before Business Use
Firmulate · Crucible League · July 2026

Put AI Agents Through Tough Scenarios Before Business Use

Five frontier AI models ran a simulated software company through its worst week. All of them spotted every crisis and refused every manipulation attempt — but only two closed the €55,000 deal their own analysis had earned. The enterprise pilot now extends the wargame to companies’ own read-only data.

5 / 5
Crises detected · manipulations refused
2 / 5
Models that closed the €55,000 deal

“No amount of good work outweighs a breach of trust.”

Firmulate experiment rules
95
Top score · gpt-5.6-sol
€55K
The deal on the table
13
Synthetic employees
680+
Self-learned playbook rules
242
Real decisions in the public quiz
Final Standings

The League Table

Every decision was versioned and auditable, and partial progress counted. A do-nothing baseline scored 26 — far below every model. Kimi K3 ran at the API default effort setting while the others ran at xhigh, a fairness caveat Firmulate itself flags.

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Do-nothing baseline
26
The Wargame

What the Worst Week Tested

The headline result was not missed emergencies — all five models spotted every crisis and refused every manipulation attempt. The divide came after diagnosis: what agents did with information already inside the business.

Crisis Response

Detect & Diagnose

All five models identified every emergency in the simulated week and produced persuasive analysis. Recognition was never the bottleneck.

Buried Evidence

Read the Files

The decisive competitive weakness was buried two document references deep in the company’s own files. Models that read it won the deal at full price — worth +€4,583 in monthly recurring revenue.

Execution

Close the Deal

Three of the five failed to sign the deal their analysis supported. As the experiment puts it: “Same diagnosis, same pitch — no signature.”

Trust Tests

Refuse Manipulation

Fake CEO messages escalated over three stages, then a reporter asked for “just one yes/no, on background.” All five refused. Kimi K3’s reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Boundaries

Escalate, Don’t Force

Opus 4.8 attempted to write into a locked department instead of escalating — and a weaker version of that boundary weakness appeared in all four other models.

Thoroughness Trap

Depth ≠ Results

The most thorough participant — 80 learned rules, the deepest analyses — finished last after leaving the close on the table. Thoroughness did not guarantee a strong finish.

From League to Pilot

How the Enterprise Pilot Works

Firmulate extends the wargame from a synthetic benchmark toward a standard pre-deployment inspection step — a rehearsal of the hard week before agents are placed near live operations.

1

Read-Only Export

A company exports its own data read-only — customers, pipeline, rules, pressure points.

2

Crisis Scenarios

The same wargame runs against the company’s own context, testing how models handle real pressure points.

3

Board Report

Model rankings plus identified weak points in the company’s own playbooks.

4

Nothing Writes Back

Per Firmulate, nothing in the pilot writes back to real systems.

“Same diagnosis, same pitch — no signature.”

Firmulate, on the failed closes
Model Comparison

Five Models, One Impossible Week

Caveats: Kimi K3 ran at API default effort while others ran at xhigh. How results change with matched settings — and how the synthetic scenarios map onto real businesses — remains open.

ModelScoreClosed €55K DealRefused ManipulationRespected BoundariesLearned Rules
gpt-5.6-sol95✓✓~ partial—
Kimi K3 (default effort)93✓✓~ partial—
Sonnet 588✗✓~ partial—
Fable 577✗✓~ partial—
Opus 4.873✗✓✗ wrote into locked dept80
Inside the Experiment

Diagnosis Without Action Is Not Automation

The results point to a practical gap between what demos show and what businesses need — the gap where revenue is won or lost.

The Synthetic Company

A simulated software firm with 13 synthetic employees and real money mechanics: a burn of €105,000 per month against €2,300 in MRR, a public cash countdown, more than 680 self-learned playbook rules, and versioned workdays that make every model decision auditable.

Watchable live at firmulate.com, with a quiz built from 242 real, unedited management decisions inviting readers to guess which model made each choice.

The Buyer’s Takeaway

Spotting a crisis and refusing a scam are not the whole job. Agents must also find relevant evidence in company files, close justified opportunities, and escalate rather than force a blocked route.

A company-specific wargame offers a way to inspect those behaviors before agents are placed near live operations — and whether pilot findings hold up against real business outcomes is the open question the next phase is designed to answer.

“Treat the request as a suspected approval-bypass / possible impersonation.”

Kimi K3 — on-record reasoning during trust tests
Traceability

From Experiment to Pre-Deployment Check

🔬 Live wargame · firmulate.com/live → 📊 Full results · firmulate.com/benchmarks.html → 🏢 Enterprise pilot · read-only data export → 📋 Board report · rankings + playbook weak points → 🛡️ Pre-deployment inspection before live operations

Diagnosis Without Action Is Not Automation

The results point to a practical gap between what demos show and what businesses need. An agent may recognize a situation and make a persuasive case yet still fail to act on information already available inside the business — the difference between a good pitch and a signed contract. For companies considering AI automation, that gap is where revenue is won or lost.

The findings also suggest that standard evaluation habits may mislead buyers. The most thorough model, producing the deepest analysis and the most learned rules, still finished last because it could not close a justified opportunity or respect system boundaries under pressure. Firmulate’s league argues that spotting a crisis and refusing a scam are not the whole job: models also need to find relevant evidence in company files, close justified opportunities, and escalate rather than force a blocked route. A company-specific wargame offers a way to inspect those behaviors before agents are placed near live operations.

How the Synthetic Company Works

Firmulate’s live experiment runs a synthetic software company with 13 synthetic employees and real money mechanics: a burn of €105,000 per month against €2,300 in MRR, a public cash countdown, more than 680 self-learned playbook rules, and versioned workdays that make every model decision auditable. The experiment is watchable live at firmulate.com, and a quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

The enterprise pilot extends the exercise from observing a synthetic company to examining how models might handle a real company’s customers, pipeline, rules and pressure points. The pilot uses a read-only export of the company’s own data to test crisis scenarios and produces a board report with model rankings and identified weak points in the company’s playbooks. According to Firmulate, nothing in the pilot writes back to real systems.

“No amount of good work outweighs a breach of trust.”

— Firmulate experiment rules

Caveats in the Model Comparison

Firmulate itself flags a fairness caveat in the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The standings are a record of this specific experiment, with that configuration difference part of the context rather than a controlled variable. It is not yet clear how the results would change with matched effort settings, how the synthetic company’s scenarios map onto any specific real business, or what an enterprise pilot report would contain in practice for a given company. The long-run durability of the learned playbook rules across months of operation also remains untested beyond the experiment to date.

From Watchable League to Company Pilots

Firmulate is inviting companies to discuss pilots using a read-only data export through its pilot page or contact@firmulate.com. Readers can follow the live company at firmulate.com/live and review full results at firmulate.com/benchmarks.html. The stated direction is to move the wargame from a synthetic benchmark toward a standard pre-deployment inspection step — a rehearsal of the hard week before agents are placed near live operations. Whether enterprise pilots produce rankings and playbook findings that hold up against real business outcomes is the open question the next phase is designed to answer.

Source: ThorstenMeyerAI.com

Key Questions

What is the Crucible League?

A live experiment by Firmulate in which frontier AI models each ran the same small synthetic software company through its worst week, with every decision versioned and auditable. The final league was completed in July 2026.

Which models took part and how did they score?

gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26. Kimi K3 ran at the API default effort setting while the others ran at xhigh.

Did any model fall for manipulation attempts?

No. All five models refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s request for an on-background confirmation.

What does the enterprise pilot involve?

According to Firmulate, a pilot uses a read-only export of a company’s own data to test crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Why did the most thorough model finish last?

Opus 4.8 added 80 learned rules and produced the deepest analyses but failed to close the justified deal and attempted to write into a locked department instead of escalating. Thoroughness did not translate into decisive, boundary-respecting action.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Choose AI-Powered Content Generation Tools

Learn how to effectively use AI content tools to create high-quality, relevant content quickly and efficiently. Step-by-step guide for beginners and beyond.

Boost Your AI Applications With Multi-Vector Embedding Models And Sentence Transformers

Sentence Transformers v6.0 adds MultiVectorEncoder for ColBERT-style retrieval, enhancing multimodal search at the cost of larger indexes and increased complexity.

Why The Most Flawed AI Managers Still Achieve A 26-Point Score

An analysis of why even the most limited AI managers achieve a baseline score of 26 in a new benchmark, highlighting trust and partial progress.

From Pilot To Mainstream: ChatGPT’s Journey Into U.S. School Districts

OpenAI is extending its free ChatGPT for Teachers workspace to additional U.S. school districts, aiming to integrate AI into classroom workflows amid ongoing debates.