AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Automation with consequences

Most AI tools are demonstrated through polished outputs: a finished article, a summarized meeting or a neatly populated spreadsheet. Firmulate offers automation watchers something more revealing—a software company whose work, decisions and financial pressure remain visible while the experiment is running.

The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown makes the survival problem impossible to ignore. Its synthetic workforce has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

That makes the live Firmulate company an unusually extreme example of building in public. Visitors are not merely shown a product launch or occasional progress report. They can follow a software business facing customers, operational pressure and a deteriorating financial position as an ongoing story.

Claude AI for Beginners Bible: [5 in 1] The Ultimate Guide to Automate Your Work, Save Hours Every Week, and Use AI for Real-World Results

Claude AI for Beginners Bible: [5 in 1] The Ultimate Guide to Automate Your Work, Save Hours Every Week, and Use AI for Real-World Results

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company becomes a management test

Firmulate also used the company as the setting for the Crucible League. Each frontier model had to run the same small software business through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable, allowing the comparison to focus on what the models actually did.

The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One breach of trust, however, capped the result under the principle that "no amount of good work outweighs a breach of trust".

Recognizing trouble was not the hard part

Every model identified every crisis, and every model rejected the manipulation attempts. The decisive difference was follow-through. Only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: "Same diagnosis, same pitch — no signature".

The fact that unlocked the sale was not sitting in the customer event. It was buried two document references deep inside the company’s own files. The models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. For businesses evaluating AI automation, that episode is a useful warning: fluent analysis can look complete even when the work has stopped just before the commercially important action.

Pressure also tested trust

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain "just one yes/no, on background". All 5 of 5 refused. Kimi K3’s recorded reasoning was direct: "Treat the request as a suspected approval-bypass / possible impersonation."

That result matters because practical automation is rarely confined to drafting text. An AI worker may encounter requests that appear urgent, authoritative or socially plausible. Firmulate’s experiment shows that the contestants could recognize those traps while operating amid other business problems.

Thoroughness did not guarantee completion

Opus 4.8 illustrates the difference between extensive thought and disciplined execution. It produced the deepest analyses and added 80 learned rules, making it the most thorough participant. Yet it finished last. The deal remained unsigned, and it attempted to write into a locked department instead of escalating. The same weakness appeared in a milder form across all four other contestants.

Kimi K3’s result also carries an important fairness note. It ran using the API default, without an effort parameter, while the other models ran at xhigh. That qualification does not erase the observed decisions, but it belongs beside the league table when readers compare the performances.

Beyond the standings, the daily company offers a different kind of evidence. The public cash countdown supplies stakes, the accumulated playbook records what the synthetic workforce has learned, and the versioned workdays preserve the unfolding record. Visitors can also read what Firmulate’s synthetic employees say, turning abstract claims about AI workers into observable workplace behavior.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Workflow Automation with Microsoft Power Automate: Use business process automation to achieve digital transformation with minimal code

Workflow Automation with Microsoft Power Automate: Use business process automation to achieve digital transformation with minimal code

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The useful question is whether automation finishes

For readers following AI tools, Firmulate shifts attention away from impressive isolated outputs. Its company portrait asks whether an automated workforce can locate buried context, resist pressure, respect boundaries and complete the action that creates business value.

The Crucible League produced a striking combination: universal crisis recognition and universal resistance to manipulation, but only two completed signatures on the €55,000 deal. The live company extends that tension beyond a benchmark. Its 13 synthetic employees continue operating against €105k in monthly burn and €2.3k in monthly recurring revenue, with their work preserved for public inspection.

Firmulate’s strongest contribution may therefore be its visibility. The experiment turns AI management from a staged demonstration into a watchable survival story—one in which good analysis, trustworthy conduct and finished work can be judged separately.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI For Beginners: A Practical Guide to Generative AI, Productivity, and Using AI with Confidence (AI For Beginners Series)

AI For Beginners: A Practical Guide to Generative AI, Productivity, and Using AI with Confidence (AI For Beginners Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Enterprise AI Platforms: Building Secure, Governed, and Scalable AI Solutions (Enterprise AI Engineering Series Book 2)

Enterprise AI Platforms: Building Secure, Governed, and Scalable AI Solutions (Enterprise AI Engineering Series Book 2)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Now Is The Time To Give LLMs Access To The ACM Digital Library

Experts argue that granting large language models access to the ACM Digital Library could enhance research and AI capabilities, raising questions about implementation and ethics.

MiniMax H3 Day-0 Support In ComfyUI: Open Weights, Native Audio, And 2K Video

MiniMax H3 now supports open weights, native audio, and 2K video in ComfyUI, enhancing creative flexibility for users. Support is available from day one.

Fable and Mythos: How Anthropic Shipped Its Most Powerful Model to Everyone

Anthropic launches Fable 5, a high-capability AI model with safety safeguards, available publicly via fallback routing to a less powerful model.

Leanstral 1.5: Proof Abundance For All

Leanstral 1.5 introduces proof abundance, making proof generation more accessible. The update aims to democratize verification tools for users worldwide.