
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Scoreboard That Doesn’t Start at Zero
If you’ve ever watched an AI tool demo and wondered whether the polished chat answers translate into actual business results, a new experiment at Firmulate offers a refreshingly blunt answer: usually not. Four frontier AI models were each handed the same job — running a small software company through its worst week, with the same customers, the same crises, and the same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The headline finding is one every automation professional should sit with: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment’s summary puts it: same diagnosis, same pitch — no signature.
As an affiliate, we earn on qualifying purchases.
So Why Does Doing Nothing Score 26?
The most interesting number in the whole exercise isn’t the winning score of 95. It’s the floor: 26. In Firmulate’s league table, a hypothetical manager that does nothing — no decisions, no actions — still walks away with 26 points. That’s not a bug; it’s the entire philosophy of the benchmark.
The reasoning is simple. In a real company, a manager who shows up, reads the situation correctly, and doesn’t make things worse has genuinely delivered partial value. A crisis that escalates unchecked is worse than a crisis that simply isn’t resolved. So partial progress counts: noticing the problem, keeping customers informed, avoiding bad decisions — these are worth real points, because in business, catastrophe avoided is a deliverable.
But the scale has a hard ceiling too. A single breach of trust caps the total grade, no matter how brilliant the rest of the performance. The benchmark’s own phrasing is worth quoting: “no amount of good work outweighs a breach of trust.” For anyone wiring AI agents into real systems, that design choice matters. It means an agent that lies once, cuts one corner on trust, cannot buy its way back with volume of output.
AI trustworthiness evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Distrust of Round Numbers
There’s a second signal embedded in the scoring: a visible skepticism of perfect 100s. The top score in the final July 2026 league table is 95, earned by gpt-5.6-sol. Behind it: Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Nobody hit 100, and the benchmark’s design suggests nobody is supposed to. A scale that routinely hands out perfect grades stops telling you anything.
business AI performance benchmarks
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Deal That Separated the Field
The decisive test wasn’t a customer shouting — it was a quiet detail. The competitor weakness that unlocked the €55,000 deal sat two document references deep in the company’s own files, not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that diagnosed the situation perfectly but never dug into their own records left the close on the table.
Then there was the social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models faced down every attempt. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
The Thoroughness Paradox
Perhaps the most instructive profile belongs to Opus 4.8: the most thorough participant in the field, generating over 80 learned rules and the deepest analyses — and finishing last. The close was left undone, and discipline slipped, including write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four models. Effort and diligence, it turns out, are not the same thing as finishing.
One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — worth keeping in mind when comparing its 93 to the field.
Watch It Happen Live
Firmulate isn’t a static report. There’s a live company running continuously: 13 synthetic employees, real money mechanics with a burn of €105k/month against €2.3k MRR, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned. You can watch it at firmulate.com/live, and the site rebuilds itself twice a day.
Want to test your own instincts? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call at firmulate.com/quiz.html. And for enterprises, the same wargame can run against a read-only export of your own business — nothing ever writes back to real systems.

The Takeaway
Most AI benchmarks measure how well a model chats. Firmulate measures how well it manages — and the gap is enormous. A scoring floor of 26 for inaction, partial credit for progress, and a hard cap for any breach of trust add up to something rare: an evaluation shaped like real accountability. For teams deploying AI automation, the lesson from the league table is uncomfortable but useful: your agent will probably spot every crisis and resist every scam. Whether it reads the file, closes the deal, and finishes what it starts — that’s where the points live.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
