📊 Full opportunity report: Discover AI’s Authentic Working Style Through This Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Firmulate has launched a live experiment testing AI models in a simulated business crisis. The test reveals differences in how models diagnose, trust, and act, providing insights into their management styles. This helps organizations evaluate AI readiness for real-world decision-making.
Firmulate.com has launched a live management experiment where five AI models are tested in a simulated business crisis, revealing their true operational styles. This experiment exposes how AI handles diagnosis, trust, and decisive action, providing critical insights for enterprise AI deployment.
The test involves five AI managers operating a small software company experiencing its worst week, with identical crises, customers, and temptations. For more on AI management styles, see the original analysis. Each model’s decisions are recorded in a transparent, auditable environment, with the goal of assessing not just analysis but also execution and follow-through.
In the final standings, gpt-5.6-sol ranked first with 95 points, followed closely by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The baseline model scored only 26, highlighting the importance of action and trust over analysis alone. Understanding these differences is part of the management test that exposes an AI’s real working style.
Despite all models recognizing crises and refusing manipulation attempts, only two signed a crucial €55,000 deal, underscoring a gap between diagnosis and execution. The models’ ability to find relevant information, escalate risks, and close deals varied significantly, revealing different management personalities and operational discipline.
Implications for AI Management and Enterprise Use
This experiment demonstrates that high-quality analysis alone does not guarantee effective management. The ability to translate diagnosis into action, trustworthiness, and follow-through is critical for AI to be operationally useful in business contexts. The findings suggest that organizations should evaluate AI models not only for their reasoning but also for their execution capabilities before deploying them at scale.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Testing and Firmulate’s Approach
Traditional AI benchmarks focus on reasoning, language understanding, or problem-solving accuracy. However, Firmulate.com has pioneered a new approach by testing AI models in a simulated business environment that mimics real-world pressure and decision-making. The live experiment involves models managing a company through crises, with decisions recorded and evaluated in a transparent setting. The July 2026 league table marks the culmination of this ongoing effort to assess AI’s practical management skills, emphasizing the importance of trust, discipline, and action.
“This experiment exposes how AI models handle real-world management tasks, revealing strengths and weaknesses that traditional benchmarks overlook.”
— Firmulate.com
business crisis simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About AI Management Performance
It is not yet clear how these findings will translate to real-world enterprise settings outside the controlled simulation. The long-term reliability of models in operational environments, especially under different pressures, remains to be tested. Additionally, the impact of varying configurations, such as API parameters, on performance is still under evaluation, as seen in the case of Kimi K3’s default settings.
enterprise AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps in AI Management Testing and Adoption
Firmulate plans to expand the experiment by incorporating more diverse scenarios and additional AI models. Enterprises are encouraged to replicate similar tests using their own business data to evaluate AI readiness before full deployment. The final league results in July 2026 will be followed by further analysis on how different management styles influence operational success, guiding future AI integration strategies.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main purpose of Firmulate’s management test?
The test aims to evaluate how AI models handle real-world management tasks, including diagnosis, trust, escalation, and execution, in a simulated crisis environment.
How are the AI models evaluated in this experiment?
Models are scored based on their decision-making, ability to find relevant information, trustworthiness, follow-through, and success in closing deals during the simulation.
What does this mean for companies considering AI automation?
Organizations should assess AI models not only for their reasoning skills but also for their operational discipline and ability to execute decisions effectively before deploying them at scale.
Are the results applicable outside the simulation?
The results provide valuable insights but need further validation in real-world, operational settings, as actual business environments may present different challenges.
What is the significance of trust and follow-through in AI decision-making?
Trust and follow-through are critical because they determine whether AI decisions lead to tangible outcomes, such as closing deals or escalating risks appropriately, which are essential for operational success.
Source: ThorstenMeyerAI.com