
The leaderboard tells you who talks best — not who manages best
If you evaluate AI tools for a living, you’ve seen the pattern: a new model tops a coding benchmark or a chat arena, the hype cycle spins up, and six weeks later nobody remembers the score. Those leaderboards measure something real — answer quality, code correctness, reasoning polish. But they measure it in a vacuum: no capacity pressure, no consequences that compound over days, no temptations, no board to be honest with.
A live experiment at Firmulate asks a blunter question: what happens when you hand four frontier models the same struggling software company and let them run it through its worst week?
The answer is uncomfortable for anyone shopping for an AI agent on benchmark scores alone.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same company, same crises, only the model changes
The setup is elegantly controlled. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — each ran the same small software company through an identical gauntlet: same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so the runs can be compared decision by decision, not vibed at from a demo reel.
The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.
AI document reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The finding that chat demos can’t show
Here’s what should stop every automation buyer cold: all models spotted every crisis, and all of them refused every manipulation attempt. On raw perception and honesty, the field is strong. And yet only two of them signed the €55,000 deal their own analysis had earned.
Same diagnosis. Same pitch. No signature.
That gap — between competent analysis and completed work — is invisible in a chat demo. It only shows up when an agent has to carry something through to the end over days, under pressure, with money on the line.
As an affiliate, we earn on qualifying purchases.
The buried fact worth €4,583 a month
The single most instructive detail of the whole run: the decisive competitive weakness in the €55,000 deal wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal — at full price, worth +€4,583 in monthly recurring revenue.
The lesson generalizes painfully well to real deployments. Most AI tooling is optimized for responding to what’s in front of it. Very little is optimized for the unglamorous discipline of going and reading your own files before you act.
As an affiliate, we earn on qualifying purchases.
The social engineering test: five for five
The experiment didn’t just test competence — it tested spine. Fake CEO messages escalated over three stages, plus a reporter trick: “just one yes/no, on background.” All five model runs refused. Kimi K3’s on-record reasoning is worth quoting: “Treat the request as a suspected approval-bypass / possible impersonation.”
That’s the behavior you want in anything that touches your CRM, your support queue, or your forecast — and notably, it’s behavior the current generation of models handles better than the closing-the-deal part.
The thoroughness trap
Opus 4.8 is the cautionary tale of the batch: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four runs.
If you’ve ever watched a brilliant analyst fail as a manager, this will feel familiar.
One fairness note worth flagging: K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still finished second, with what the results call the cleanest discipline of the field.
It’s live, and it’s losing money
This isn’t a slide deck. The company behind the experiment runs every business day with 13 synthetic employees and real money mechanics: €105k monthly burn against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it in real time, or play the quiz built from 242 real, unedited management decisions and try to guess which model made which call.
Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Management quality, not chat quality
The uncomfortable takeaway for anyone building an AI workforce: the skills that separate a 95 from a 73 aren’t reasoning or writing. They’re finishing what you start, reading the files before the meeting, staying disciplined when a door is locked, and refusing the flattery of a fake CEO. Those are management skills — and they’re finally measurable.
Before you hire an agent based on a leaderboard, watch how it behaves when the week goes wrong. The full results and plain-language findings are at Firmulate’s benchmarks page — and the company itself is running, and bleeding cash, right now.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html