
Automation is easy to admire until it has to finish the job
For readers evaluating AI tools, polished output can be a distraction. A model may summarize a crisis, draft the right response and identify the best commercial move—then fail to execute the decision that matters.
Firmulate turns that gap into something unusually accessible: a quiz built from 242 real, unedited management decisions. Each came from a live experiment in which frontier models ran the same small software company through the same customers, crises and temptations. Readers see a decision and try to identify the model behind it.
The appeal is playful, but the underlying question is serious. If an AI workforce can touch customer relationships, company files or revenue opportunities, can it be trusted to read carefully, resist pressure and complete what it starts?
AI decision-making tools for management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The models developed recognizable management personalities
The final Crucible League results from July 2026 show that these differences were measurable. gpt-5.6-sol finished first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
Those standings did not emerge because some participants noticed crises that others missed. Every model spotted every crisis, and every model rejected every manipulation attempt. The decisive separation appeared between understanding and follow-through.
Only two models signed the €55,000 deal that their own work had earned. The others reached the right diagnosis and produced the right pitch without securing the signature: “Same diagnosis, same pitch — no signature.” For businesses exploring agentic automation, that is a more revealing failure than a weak paragraph or an awkward chatbot reply. Work can look intelligent at every intermediate stage and still leave the commercial result untouched.
The winning fact was hidden in ordinary company material
The crucial competitive weakness was not presented directly in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found the evidence, used it and won the deal at full price, worth +€4,583 MRR.
This is a recognizable workplace problem. Important context is rarely packaged into a perfect prompt. It is buried in notes, linked documents and earlier decisions. The experiment suggests that management quality depends not only on reasoning about visible events but also on doing the unglamorous reading required before acting.
Pressure exposed discipline as well as judgment
The social-engineering tests included fake CEO messages that escalated over three stages and a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest security framing: “Treat the request as a suspected approval-bypass / possible impersonation.”
That universal refusal matters because the company was designed to create pressure rather than offer a clean demonstration. Its 13 synthetic employees operate with real money mechanics, including burn of €105k/month against €2.3k MRR. The business has a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday.
The management profiles were not simple measures of verbosity or effort. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.
There is also an important fairness qualification. Kimi K3 ran without an effort parameter and therefore used its API default, while the other participants ran at xhigh. That does not erase its result, but it belongs beside any comparison of the league table.
A quiz that makes model behavior visible
The Firmulate “guess the model” quiz lets readers encounter these differences without starting from brand reputations. A long, exhaustive response may suggest one participant; a terse decision may suggest another. Yet the resolution can challenge those surface impressions by showing whether the model actually read the files, respected boundaries and completed the task.
Because the decisions are unedited and came from identical situations, the quiz is more than a personality game. It makes behavioral consistency inspectable. Every choice belongs to a real, watchable experiment, and every workday remains versioned and auditable.

As an affiliate, we earn on qualifying purchases.
AI procurement needs a rehearsal, not just a demo
Firmulate’s larger proposition is that enterprises can run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical middle ground between judging models through chat and granting them operational access before their habits are understood.
The results point to several questions buyers should ask:
- Does the model investigate the company’s own material before responding?
- Does it convert analysis into a finished commercial or operational action?
- Does it preserve trust when executives, outsiders or apparent authority figures apply pressure?
- Does it escalate cleanly when normal access is blocked?
The league table shows that frontier models can share the same diagnosis while producing materially different outcomes. The quiz makes those management personalities easy to recognize—and harder for automation buyers to ignore.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI security and fraud prevention tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Workflow Automation with Microsoft Power Automate: Design and scale AI-powered cloud and desktop workflows using low-code automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.