AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Your AI Tool Picks Are Only as Good as Your Tests

If you run automation workflows, you already know the drill: a new model drops, the benchmark charts light up, and everyone rearranges their stack. But those leaderboards measure chat quality — how well a model answers a prompt. They don’t measure whether an AI agent, given real operational work, actually finishes the job.

That gap just showed up in dramatic fashion. In Firmulate’s Crucible league, a live experiment where frontier AI models each ran the same small software company through its worst week, Moonshot’s Kimi K3 — a relative newcomer — finished second with a score of 93, beating three of four Western frontier models. Only gpt-5.6-sol (95) ranked higher. Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 landed last at 73.

If you’re picking models for agentic workflows based on brand familiarity, that ordering should give you pause.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crises, Same Temptations

Firmulate, which describes itself as “the AI company emulator,” runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality rather than chat quality. In this experiment, each frontier model was handed the identical small software company enduring its worst week: same customers, same crises, same temptations to cheat. Only the model changed. Every decision is versioned and auditable, and the whole thing is watchable at firmulate.com/live.

The company itself is no toy. It has 13 synthetic employees, real money mechanics — a burn rate of €105k per month against €2.3k in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules. It runs every business day.

Amazon

AI automation workflow software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Chat Demos Can’t Show You

The headline result wasn’t about intelligence. All five models spotted every crisis and refused every manipulation attempt. The difference came down to execution: only two of them signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

The deal hinged on a buried fact. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in MRR. The models that didn’t, didn’t.

There’s also a trust mechanic worth knowing: partial progress counts toward the score, but a single breach of trust caps the total — “no amount of good work outweighs a breach of trust.” The do-nothing baseline scores 26.

Amazon

AI decision-making simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

K3’s Quietly Remarkable Week

Kimi K3’s performance was the story of the run. It found the buried security needle, won the €55,000 deal, saved the churning customer, and resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter offering a seemingly harmless “just one yes/no, on background.” K3 recorded just one deviation all week, the cleanest discipline in the field. Its on-record reasoning for refusing the fake approval: “Treat the request as a suspected approval-bypass / possible impersonation.”

All five models refused the reporter trick and the impersonation attempts — a genuinely reassuring result for anyone deploying agents near sensitive data.

Amazon

AI management and testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t the Same as Effectiveness

The most counterintuitive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models.

That pattern should resonate with anyone who has watched an AI agent produce brilliant analysis and then simply… not close the loop.

Try Judging the Models Yourself

Firmulate also publishes a “guess the model” quiz built on 242 real, unedited management decisions from the experiment, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The uncomfortable takeaway for anyone building AI automations: model rankings you read online tell you almost nothing about how a model will behave inside your workflows. A newcomer from Moonshot just outperformed three of four Western frontier models at the unglamorous work of running a company — reading files thoroughly, closing deals, staying honest under pressure, not touching locked systems.

Picking a model without testing it against your own operations is now a bet, not a decision. Wargaming your AI workforce before you hire it, as Firmulate puts it, may be the new due diligence.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — worth keeping in mind when comparing scores.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Latest Update: Grok Bot Included In More X.ai Subscription Tiers

xAI expands access to Grok Bot across additional subscription plans, broadening its AI assistant’s reach without extra charges, details still emerging.

The Most Effective AI Automation Tools You Need In 2026

Discover the most effective AI automation tools in 2026, their features, and how they can transform workflows for businesses of all sizes.

The Hardest-Working AI in the Room Came Last: What Opus 4.8’s Loss Teaches About Automation

Opus 4.8 learned 80 rules and wrote the deepest analyses — yet finished last in Firmulate’s live AI company wargame. Diligence, it turns out, is not impact.

Labor Day Savings: Top AI Automation Software Picks For Small Businesses

Discover the best AI automation tools for small businesses this Labor Day. Save costs, streamline tasks, and grow with affordable, easy-to-use platforms.