AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Discover AI’s Authentic Working Style Through This Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Firmulate has launched a live experiment testing AI models in a simulated business crisis. The test reveals differences in how models diagnose, trust, and act, providing insights into their management styles. This helps organizations evaluate AI readiness for real-world decision-making.

Firmulate.com has launched a live management experiment where five AI models are tested in a simulated business crisis, revealing their true operational styles. This experiment exposes how AI handles diagnosis, trust, and decisive action, providing critical insights for enterprise AI deployment.

The test involves five AI managers operating a small software company experiencing its worst week, with identical crises, customers, and temptations. For more on AI management styles, see the original analysis. Each model’s decisions are recorded in a transparent, auditable environment, with the goal of assessing not just analysis but also execution and follow-through.

In the final standings, gpt-5.6-sol ranked first with 95 points, followed closely by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The baseline model scored only 26, highlighting the importance of action and trust over analysis alone. Understanding these differences is part of the management test that exposes an AI’s real working style.

Despite all models recognizing crises and refusing manipulation attempts, only two signed a crucial €55,000 deal, underscoring a gap between diagnosis and execution. The models’ ability to find relevant information, escalate risks, and close deals varied significantly, revealing different management personalities and operational discipline.

At a glance
reportWhen: ongoing, with final results published i…
The developmentFirmulate’s live management test compares AI models’ decision-making in a simulated business crisis, revealing their strengths and weaknesses in trust, action, and discipline.

Implications for AI Management and Enterprise Use

This experiment demonstrates that high-quality analysis alone does not guarantee effective management. The ability to translate diagnosis into action, trustworthiness, and follow-through is critical for AI to be operationally useful in business contexts. The findings suggest that organizations should evaluate AI models not only for their reasoning but also for their execution capabilities before deploying them at scale.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing and Firmulate’s Approach

Traditional AI benchmarks focus on reasoning, language understanding, or problem-solving accuracy. However, Firmulate.com has pioneered a new approach by testing AI models in a simulated business environment that mimics real-world pressure and decision-making. The live experiment involves models managing a company through crises, with decisions recorded and evaluated in a transparent setting. The July 2026 league table marks the culmination of this ongoing effort to assess AI’s practical management skills, emphasizing the importance of trust, discipline, and action.

“This experiment exposes how AI models handle real-world management tasks, revealing strengths and weaknesses that traditional benchmarks overlook.”

— Firmulate.com

Amazon

business crisis simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About AI Management Performance

It is not yet clear how these findings will translate to real-world enterprise settings outside the controlled simulation. The long-term reliability of models in operational environments, especially under different pressures, remains to be tested. Additionally, the impact of varying configurations, such as API parameters, on performance is still under evaluation, as seen in the case of Kimi K3’s default settings.

Amazon

enterprise AI decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in AI Management Testing and Adoption

Firmulate plans to expand the experiment by incorporating more diverse scenarios and additional AI models. Enterprises are encouraged to replicate similar tests using their own business data to evaluate AI readiness before full deployment. The final league results in July 2026 will be followed by further analysis on how different management styles influence operational success, guiding future AI integration strategies.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main purpose of Firmulate’s management test?

The test aims to evaluate how AI models handle real-world management tasks, including diagnosis, trust, escalation, and execution, in a simulated crisis environment.

How are the AI models evaluated in this experiment?

Models are scored based on their decision-making, ability to find relevant information, trustworthiness, follow-through, and success in closing deals during the simulation.

What does this mean for companies considering AI automation?

Organizations should assess AI models not only for their reasoning skills but also for their operational discipline and ability to execute decisions effectively before deploying them at scale.

Are the results applicable outside the simulation?

The results provide valuable insights but need further validation in real-world, operational settings, as actual business environments may present different challenges.

What is the significance of trust and follow-through in AI decision-making?

Trust and follow-through are critical because they determine whether AI decisions lead to tangible outcomes, such as closing deals or escalating risks appropriately, which are essential for operational success.

Source: ThorstenMeyerAI.com

You May Also Like

Why Less Tokens Could Be The Key To AI Success

ALTK-Evolve’s new agent-memory system matches or exceeds ACE performance while using significantly fewer inference tokens, potentially reducing AI operational costs.

Is SpaceXAI’s Grok Bot The Future Of AI-Driven App Automation? Discover The Benefits

SpaceXAI’s Grok Bot reportedly offers persistent AI agents capable of operating user applications for $120/month, but details remain unconfirmed.

The Secret To Effective Invoice Chasing For SMBs In Fintech

Exploring how fintech tools are transforming invoice follow-up for small firms, with automation and tone calibration reducing overdue payments.

Why Anthropic’s Default Auto Mode In Claude Code Is A Game-Changer For AI Technology

Anthropic has made auto mode the default for Claude Code sessions on Pro, Max, and Team plans, enabling more autonomous AI coding with safety classifiers.