AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

An AI agent can handle the workflow—and still leave the deal unsigned

For businesses exploring AI tools and automation, a smooth demo can make a system look ready for real work. But a model that spots a sales opportunity, explains why it matters and drafts the right pitch may still fail at the final step. Firmulate’s live company experiment puts that gap on display: the work is not just to recognize a crisis or recommend an action, but to carry it through.

That distinction matters when automation touches customers, revenue and trust. Firmulate’s public experiment follows AI models running the same small software company through its worst week. The decisions are real within the experiment and watchable at Firmulate.

Same company, same crises, different outcomes

In the final Crucible League, dated July 2026, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The striking result was not that the models missed danger. Every model spotted every crisis and refused every manipulation attempt. The difference came at the close: only two signed the €55,000 deal their own analysis had earned. The experiment’s summary is pointed: “Same diagnosis, same pitch — no signature.” For anyone automating sales or operations, recognition and recommendation are not the same as execution.

The useful clue was buried in the company’s own files

The deal hinged on a competitor weakness hidden two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that useful business context may sit outside the immediate conversation or record an agent is handling.

Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Those refusals show sound instincts under pressure; the deal results show that integrity alone does not guarantee follow-through.

More effort did not guarantee a better finish

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and slipped in discipline, attempting writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is also a fairness detail in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The leaderboard is a useful account of this experiment, not a promise that any model will behave the same way in every company.

From watching an experiment to testing your own business

The live company has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call. The experiment offers a view of how models respond when choices accumulate across a business day, rather than in a one-off chat.

For an enterprise, the next step is a pilot against its own business. Firmulate can use a read-only export to create a digital twin, run crisis scenarios against the company’s own context, and produce a board report with model rankings and weak points in existing playbooks. Nothing writes back to real systems. That gives decision-makers a way to examine how an AI workforce might respond before placing it in live operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the judgment before you hand over the workflow

Firmulate’s experiment suggests that an AI agent can identify the problem, resist manipulation and still fall short of completing the work. A company-specific wargame makes that gap visible against your own customers, procedures and pressure points. To discuss an enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Small Models Have Arrived

Small AI models are now available, offering efficient alternatives to large models. This development impacts AI accessibility and deployment.

Discover AI’s Authentic Working Style Through This Management Test

Firmulate’s management challenge reveals how AI models handle crisis, trust, and action in a simulated business environment, exposing their true operational capabilities.

AI Innovation In Progress: Anthropic’s Claude And The Next Version Of Itself

Anthropic states its AI model, Claude, is actively used to help build its successor, marking a step toward AI-assisted model development, though details remain unverified.

The Turbulent AI Era Is Here

The AI sector is experiencing significant upheaval due to rapid technological developments, regulatory debates, and ethical concerns, marking a turbulent era for AI.