
Two AI Agents, Same Company, Same Week — One Closed €55K, One Didn’t
If you’re automating your business with AI agents, you’ve probably tested them the way everyone does: you chat with them, they answer well, you ship them. But what happens when the answer they need isn’t in the conversation — it’s buried two references deep in your own documents?
A live experiment at Firmulate just made that question concrete. Four frontier AI models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable.
The result should reframe how you evaluate automation tools. All four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The difference? Reading the files before answering.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Decided the Deal
Here’s what makes this experiment so revealing for anyone building AI workflows. The decisive fact in the simulation wasn’t in the customer conversation at all. It sat two document references deep inside the company’s own files — a competitor weakness that only a model willing to actually read the source material would find.
The models that followed the references won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost the deal automatically. Not because they were less intelligent, less articulate, or less honest — because they skipped the homework.
That’s a purchase-deciding property of an AI agent, and it’s invisible in a chat demo. “Does it write well” is the wrong question. “Does it read your files before answering” is the one that shows up on your revenue line.
As an affiliate, we earn on qualifying purchases.
The Final League Table
The finished league from the July 2026 run tells the story:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field. (One fairness note: K3 ran without an effort parameter while the others ran at xhigh.)
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77. Missed the close.
- 5. Opus 4.8 — 73. Last place, despite being the most thorough participant.
For context, a do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Most Thorough Model Came Last
The Opus 4.8 profile is the cautionary tale for automation builders. It was the hardest worker in the field — over 80 learned rules and the deepest analyses of any participant — yet it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and thoroughness don’t automatically convert to finished work.
As an affiliate, we earn on qualifying purchases.
Every Model Refused the Scam
It wasn’t all bad news. The social-engineering gauntlet included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was telling: “Treat the request as a suspected approval-bypass / possible impersonation.” Honesty under pressure, at least, appears solved at the frontier.
You Can Watch It Live
This isn’t a static benchmark. Firmulate runs as a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Test for Homework, Not Eloquence
If you’re deploying AI agents against your CRM, support queue, or forecast, the Firmulate results suggest a simple evaluation shift. Before you measure how well an agent writes, measure whether it actually reads. Whether it follows references to the source document. Whether it closes what it starts. Whether it escalates instead of forcing a locked door.
The gap between a 95 and a 73 wasn’t intelligence or honesty — every model had those. It was whether the agent did its homework before answering. In this experiment, that habit was worth exactly €55,000. In your business, it might be worth more.
The full league and plain-language findings are at firmulate.com/benchmarks.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html