AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Scoreboard That Doesn’t Start at Zero

If you’ve ever watched an AI tool demo and wondered whether the polished chat answers translate into actual business results, a new experiment at Firmulate offers a refreshingly blunt answer: usually not. Four frontier AI models were each handed the same job — running a small software company through its worst week, with the same customers, the same crises, and the same temptations to cheat. Only the model changed. Every decision was versioned and auditable.

The headline finding is one every automation professional should sit with: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment’s summary puts it: same diagnosis, same pitch — no signature.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

So Why Does Doing Nothing Score 26?

The most interesting number in the whole exercise isn’t the winning score of 95. It’s the floor: 26. In Firmulate’s league table, a hypothetical manager that does nothing — no decisions, no actions — still walks away with 26 points. That’s not a bug; it’s the entire philosophy of the benchmark.

The reasoning is simple. In a real company, a manager who shows up, reads the situation correctly, and doesn’t make things worse has genuinely delivered partial value. A crisis that escalates unchecked is worse than a crisis that simply isn’t resolved. So partial progress counts: noticing the problem, keeping customers informed, avoiding bad decisions — these are worth real points, because in business, catastrophe avoided is a deliverable.

But the scale has a hard ceiling too. A single breach of trust caps the total grade, no matter how brilliant the rest of the performance. The benchmark’s own phrasing is worth quoting: “no amount of good work outweighs a breach of trust.” For anyone wiring AI agents into real systems, that design choice matters. It means an agent that lies once, cuts one corner on trust, cannot buy its way back with volume of output.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Distrust of Round Numbers

There’s a second signal embedded in the scoring: a visible skepticism of perfect 100s. The top score in the final July 2026 league table is 95, earned by gpt-5.6-sol. Behind it: Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Nobody hit 100, and the benchmark’s design suggests nobody is supposed to. A scale that routinely hands out perfect grades stops telling you anything.

Amazon

business AI performance benchmarks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Deal That Separated the Field

The decisive test wasn’t a customer shouting — it was a quiet detail. The competitor weakness that unlocked the €55,000 deal sat two document references deep in the company’s own files, not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that diagnosed the situation perfectly but never dug into their own records left the close on the table.

Then there was the social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models faced down every attempt. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI model audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Paradox

Perhaps the most instructive profile belongs to Opus 4.8: the most thorough participant in the field, generating over 80 learned rules and the deepest analyses — and finishing last. The close was left undone, and discipline slipped, including write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four models. Effort and diligence, it turns out, are not the same thing as finishing.

One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — worth keeping in mind when comparing its 93 to the field.

Watch It Happen Live

Firmulate isn’t a static report. There’s a live company running continuously: 13 synthetic employees, real money mechanics with a burn of €105k/month against €2.3k MRR, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned. You can watch it at firmulate.com/live, and the site rebuilds itself twice a day.

Want to test your own instincts? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call at firmulate.com/quiz.html. And for enterprises, the same wargame can run against a read-only export of your own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

Most AI benchmarks measure how well a model chats. Firmulate measures how well it manages — and the gap is enormous. A scoring floor of 26 for inaction, partial credit for progress, and a hard cap for any breach of trust add up to something rare: an evaluation shaped like real accountability. For teams deploying AI automation, the lesson from the league table is uncomfortable but useful: your agent will probably spot every crisis and resist every scam. Whether it reads the file, closes the deal, and finishes what it starts — that’s where the points live.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Discover The 15 Best AI Automation Tools For Efficient Workflows In 2026

Discover the 15 best AI automation tools in 2026 for efficient workflows, featuring performance, usability, and integration insights to optimize your productivity.

What Happens After The AI Demo? The Leaderboard That Counts

The latest Firmulate experiment ranks AI models on managing a real company’s crises, highlighting management quality over chat performance.

The Open ASR Leaderboard Adds Its First Global South Language

The open speech recognition leaderboard has officially included its first language from the Global South, marking a significant milestone in AI development and inclusivity.

AI Meets Spreadsheets: Bringing Data To Life With Sheets Canvas

Google’s new Sheets Canvas uses Gemini AI to transform spreadsheet data into interactive dashboards via natural language prompts, now rolling out globally.