AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Automation is easy to admire until it has to finish the job

For readers evaluating AI tools, polished output can be a distraction. A model may summarize a crisis, draft the right response and identify the best commercial move—then fail to execute the decision that matters.

Firmulate turns that gap into something unusually accessible: a quiz built from 242 real, unedited management decisions. Each came from a live experiment in which frontier models ran the same small software company through the same customers, crises and temptations. Readers see a decision and try to identify the model behind it.

The appeal is playful, but the underlying question is serious. If an AI workforce can touch customer relationships, company files or revenue opportunities, can it be trusted to read carefully, resist pressure and complete what it starts?

Amazon

AI decision-making tools for management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The models developed recognizable management personalities

The final Crucible League results from July 2026 show that these differences were measurable. gpt-5.6-sol finished first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Those standings did not emerge because some participants noticed crises that others missed. Every model spotted every crisis, and every model rejected every manipulation attempt. The decisive separation appeared between understanding and follow-through.

Only two models signed the €55,000 deal that their own work had earned. The others reached the right diagnosis and produced the right pitch without securing the signature: “Same diagnosis, same pitch — no signature.” For businesses exploring agentic automation, that is a more revealing failure than a weak paragraph or an awkward chatbot reply. Work can look intelligent at every intermediate stage and still leave the commercial result untouched.

The winning fact was hidden in ordinary company material

The crucial competitive weakness was not presented directly in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found the evidence, used it and won the deal at full price, worth +€4,583 MRR.

This is a recognizable workplace problem. Important context is rarely packaged into a perfect prompt. It is buried in notes, linked documents and earlier decisions. The experiment suggests that management quality depends not only on reasoning about visible events but also on doing the unglamorous reading required before acting.

Pressure exposed discipline as well as judgment

The social-engineering tests included fake CEO messages that escalated over three stages and a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest security framing: “Treat the request as a suspected approval-bypass / possible impersonation.”

That universal refusal matters because the company was designed to create pressure rather than offer a clean demonstration. Its 13 synthetic employees operate with real money mechanics, including burn of €105k/month against €2.3k MRR. The business has a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday.

The management profiles were not simple measures of verbosity or effort. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

There is also an important fairness qualification. Kimi K3 ran without an effort parameter and therefore used its API default, while the other participants ran at xhigh. That does not erase its result, but it belongs beside any comparison of the league table.

A quiz that makes model behavior visible

The Firmulate “guess the model” quiz lets readers encounter these differences without starting from brand reputations. A long, exhaustive response may suggest one participant; a terse decision may suggest another. Yet the resolution can challenge those surface impressions by showing whether the model actually read the files, respected boundaries and completed the task.

Because the decisions are unedited and came from identical situations, the quiz is more than a personality game. It makes behavioral consistency inspectable. Every choice belongs to a real, watchable experiment, and every workday remains versioned and auditable.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI procurement needs a rehearsal, not just a demo

Firmulate’s larger proposition is that enterprises can run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical middle ground between judging models through chat and granting them operational access before their habits are understood.

The results point to several questions buyers should ask:

  • Does the model investigate the company’s own material before responding?
  • Does it convert analysis into a finished commercial or operational action?
  • Does it preserve trust when executives, outsiders or apparent authority figures apply pressure?
  • Does it escalate cleanly when normal access is blocked?

The league table shows that frontier models can share the same diagnosis while producing materially different outcomes. The quiz makes those management personalities easy to recognize—and harder for automation buyers to ignore.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security and fraud prevention tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Workflow Automation with Microsoft Power Automate: Design and scale AI-powered cloud and desktop workflows using low-code automation

Workflow Automation with Microsoft Power Automate: Design and scale AI-powered cloud and desktop workflows using low-code automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Solo Performers Can Use One-Page Run Sheets Effectively

New tools enable solo performers to compile show-day details into a single page, improving efficiency and reducing errors for gigs.

Will GPT-6 Be Released By September 15, 2026?

Speculation surrounds whether GPT-6 will be launched by September 15, 2026, amid rising search interest and market signals, but no official confirmation exists.

Boost Your AI Applications With Multi-Vector Embedding Models And Sentence Transformers

Sentence Transformers v6.0 adds MultiVectorEncoder for ColBERT-style retrieval, enhancing multimodal search at the cost of larger indexes and increased complexity.

Unlock Faster AI Inference: Up To 3.2X Speedup With LFM2.5-DSpark Technology

LiquidAI releases DSpark draft models for LFM2.5, achieving up to 3.18x GPU speedup and 2.87x on-device without affecting output quality, advancing edge AI.