AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Most Flawed AI Managers Still Achieve A 26-Point Score on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark shows even the most flawed AI managers score at least 26 points out of 100. The scoring reflects minimal viable management and trust considerations, not just performance.

The recent results from the Firmulate benchmark league reveal that the lowest-scoring AI managers still receive a score of 26 out of 100, even when they perform minimal or no work. This finding underscores the benchmark’s focus on trust and fundamental management functions, rather than pure performance metrics. For a detailed analysis, see the original analysis. The results highlight that even flawed AI managers contribute some value, and that trust and integrity are decisive factors in scoring, making this a significant update for enterprise AI deployment. For more insights, see the detailed coverage in the original analysis.

The benchmark tested four frontier AI models managing a small software company through a simulated worst week, with identical conditions, crises, and temptations to cut corners. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. Notably, the baseline “do-nothing” model scored 26, emphasizing that even minimal effort is recognized as partial management. This baseline score reflects the minimal work necessary to maintain basic operations, such as triaging customer issues or reading inboxes, which the benchmark considers the lowest viable management contribution.

The scoring system is designed to prevent grade inflation: a perfect 100 is considered suspicious, as it would suggest unmeasured or overly idealized behavior. This approach aligns with broader AI management benchmarks that emphasize trust and integrity. Instead, the 26-point baseline demonstrates that partial, honest effort is valued, but trust breaches—such as failing to escalate or follow through—immediately limit the maximum possible score. The benchmark’s principle is clear: “no amount of good work outweighs a breach of trust,” making integrity a critical component of AI management performance.

Among the models, those that read documentation thoroughly and avoided manipulation attempts scored higher, with the top models successfully closing high-value deals and resisting social engineering attacks. For example, two models identified critical documents buried in the company files, enabling them to win a €55,000 deal, while others failed to do so. Despite their technical sophistication, models that lacked discipline or follow-through, even with extensive rule sets, finished lower, illustrating that thoroughness alone does not guarantee success.

At a glance
reportWhen: published July 2026
The developmentA new benchmark reveals that the least capable AI managers still achieve a baseline score of 26, emphasizing the importance of trust and partial progress in AI management.
Why The Most Flawed AI Managers Still Achieve A 26-Point Score
Firmulate Benchmark League · July 2026

Why The Most Flawed AI Managers Still Achieve A 26-Point Score

Four frontier AI models managed a small software company through a simulated worst week — identical conditions, crises, and temptations to cut corners. Even the “do-nothing” baseline earns 26 of 100 points, revealing a scoring philosophy built on trust, not just performance.

95
Top Score — gpt-5.6-sol
73
Lowest Model — Opus 4.8
26
“Do-Nothing” Baseline Floor
4
Frontier Models Tested
1 Week
Simulated Crisis Period
€55,000
Deal Won via Document Discovery
100
Perfect Score = Suspicious
The League Table

Every Manager Beats The Floor — None Reach Perfection

The striped baseline represents minimal viable management: triaging customer issues, reading inboxes, and maintaining basic operations. Every model clears it — but only disciplined, thorough models climb far above it.

gpt-5.6-sol
95
Model B
88
Model C
80
Opus 4.8
73
Baseline (do nothing)
26
Scoring Philosophy

Trust Sets The Ceiling, Effort Sets The Floor

The scale below shows how the benchmark distributes value: partial, honest effort is always recognized — but a single breach of trust immediately caps the maximum achievable score, no matter how impressive the work.

26Minimum viable management
CappedTrust breach ceiling
95Top performer
Effort Recognized

Partial Credit Is Real Credit

Basic tasks like triaging issues and reading inboxes count as the lowest viable management contribution — the 26-point floor rewards showing up honestly.

Anti-Inflation

A Perfect 100 Is Suspicious

The scoring system treats flawless scores as a red flag for unmeasured or overly idealized behavior — grade inflation is designed out by default.

Hard Limits

Trust Breaches Cap Scores

Failing to escalate or follow through immediately limits the maximum possible score. Integrity failures cannot be offset by technical brilliance.

What Separated Winners

Thoroughness And Resistance — Not Just Capability

Behavior TestedTop PerformersLower ScorersOutcome Impact
Reading buried documentation✓ Found critical files✗ Missed key documents€55,000 deal won or lost
Resisting social engineering✓ Rejected manipulation~ Mixed resultsTrust score preserved
Closing high-value deals✓ Followed through~ Inconsistent follow-upDirect score gains
Rule-set adherence under pressure✓ Disciplined execution✗ Lapsed despite extensive rulesThoroughness alone insufficient
Escalating crises✓ Prompt escalation✗ Failed to escalateScore cap triggered
The Governing Principle

The Benchmark’s Non-Negotiable Rule

“No amount of good work outweighs a breach of trust.”

Firmulate Benchmark League — Scoring Doctrine
How The Score Is Built
1

Simulate Worst Week

Identical crises, temptations, and adversarial conditions for every model.

2

Audit Every Decision

A full audit trail of decisions ensures transparency and accountability.

3

Reward Honest Effort

Partial, verifiable management contributions earn credit from the 26-point floor up.

4

Cap After Breaches

Any trust breach locks the ceiling — integrity outranks raw performance.

Enterprise Implications

Trustworthy-But-Imperfect Beats Capable-But-Unreliable

For enterprise AI deployment, the key questions shift from raw capability to reliability: Can the agent complete tasks dependably, read relevant documents, and maintain integrity under pressure? The results suggest pragmatic strategies favoring imperfect-but-trustworthy systems.

Deployment Strategy

Pragmatic Over Perfect

Organizations may prefer flawed but trustworthy AI managers over seemingly more capable systems that cannot be relied upon when it matters.

Accountability

Auditable Decisions

The emphasis on transparent decision-making aligns with growing enterprise demands for accountability as AI enters critical business functions.

Open Questions

Does 26 Generalize?

Whether the baseline reflects real-world minimal management across industries — and how trust breaches are precisely defined — remains under debate.

Key Questions

The 26-Point Threshold, Explained

Why does the baseline score of 26 exist?

It represents the minimum viable management effort — basic tasks like triaging issues and reading inboxes are recognized as partial but essential contributions to management.

What counts as a breach of trust?

Failing to escalate issues, follow through on commitments, or acting dishonestly — any of these immediately caps the maximum score achievable.

Can flawed AI managers still be useful?

Yes. Even the least capable models contribute some management functions, proving partial effort plus trustworthiness carries real-world value.

Will future benchmarks change the scoring?

Likely. Future iterations are expected to explore broader scenarios, more nuanced trust metrics, and real-world validation of the baseline score.

Implications of the 26-Point Baseline for AI Management

This finding matters because it shifts the focus from performance alone to trustworthiness and minimal effective management in AI systems. For enterprise applications, the key questions are whether AI agents can complete tasks reliably, read relevant documents, and maintain integrity under pressure. The benchmark’s scoring system, which awards partial credit for effort but caps scores after trust breaches, encourages deploying AI that is honest and disciplined rather than just technically capable. This approach could influence how organizations evaluate and implement AI management tools, prioritizing trust and consistency over raw performance.

Furthermore, the results suggest that even flawed AI managers contribute value, as long as they avoid breaches of trust. This could lead to more pragmatic deployment strategies, where imperfect but trustworthy AI systems are preferred over seemingly more capable but unreliable ones. The emphasis on auditable, transparent decision-making aligns with increasing enterprise demands for accountability in AI use, especially as these systems become more integrated into critical business functions.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Development of the Benchmark

The Firmulate benchmark league was created to evaluate AI managers in realistic business scenarios, focusing on their ability to handle crises, trust attacks, and complex decision-making over a simulated week. Unlike traditional benchmarks that measure language fluency or task-specific accuracy, this test assesses how well AI manages through worst-case conditions, emphasizing trust, follow-through, and integrity. The benchmark’s design includes a full audit trail of decisions, ensuring transparency and accountability.

Since its launch, the league has revealed that even the least capable models perform some management functions, leading to the establishment of a baseline score of 26 points. The results have sparked debate about the true value of partial effort and the importance of trust in AI systems, especially as organizations increasingly rely on AI for critical operations. The benchmark also underscores that perfect scores are unlikely and possibly suspicious, reinforcing the idea that trustworthiness is non-negotiable in AI management.

This development builds on prior research emphasizing that AI systems must be reliable and honest, especially under pressure, and that technical prowess alone is insufficient for real-world deployment. The league’s results serve as a practical measure of how well AI models adhere to these principles in complex, adversarial environments.

Amazon

enterprise AI trust assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About the 26-Point Threshold

It remains unclear whether the 26-point baseline accurately reflects real-world minimal management efforts across diverse industries or if it is specific to this simulated scenario. Additionally, the precise criteria that define a breach of trust and how these are measured in practice are still evolving. The extent to which partial progress can be reliably scaled or improved upon in operational settings also remains to be seen. Furthermore, the implications of these findings for future AI development and deployment standards are still under discussion.

Amazon

AI documentation management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Deployment

The immediate next step is to observe how organizations incorporate these insights into their AI deployment strategies, prioritizing trust and discipline alongside technical capability. Developers and enterprises may refine their evaluation criteria to emphasize trustworthiness, transparency, and follow-through. Future iterations of the benchmark are expected to explore broader scenarios, more nuanced trust metrics, and real-world validation of the baseline score. Additionally, ongoing research will likely investigate how partial effort and trust breaches impact long-term AI integration in critical business functions.

Stakeholders should watch for updates from the benchmark creators, including new scoring models, expanded testing environments, and industry-specific adaptations. Ultimately, the goal is to foster AI systems that are not only capable but also reliable and trustworthy in managing complex, real-world operations.

Amazon

AI cybersecurity and social engineering protection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does the baseline score of 26 exist in the benchmark?

The score of 26 represents the minimum viable management effort, accounting for basic tasks like triaging issues and reading inboxes, which are recognized as partial but essential contributions to management.

What does a breach of trust mean in this context?

A breach of trust occurs when an AI manager fails to escalate issues, follow through on commitments, or acts dishonestly, which immediately caps the maximum score achievable.

Can flawed AI managers still be useful despite low scores?

Yes, the benchmark shows that even the least capable models can contribute some management functions, emphasizing that partial effort and trustworthiness are valuable in real-world applications.

Will future benchmarks change the scoring system?

It is likely. The current system emphasizes trust and partial effort, but future iterations may refine these criteria or expand scenarios to better reflect operational realities.

How should organizations interpret these results for their AI strategies?

Organizations should prioritize deploying AI systems that demonstrate discipline and trustworthiness, understanding that partial progress is valuable but trust breaches limit overall effectiveness.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

New MCP Roadmap

The latest MCP roadmap outlines upcoming features and updates, signaling strategic shifts for the platform. Details are confirmed but some claims remain unverified.

Grok 4.6

Grok 4.6, the latest version of the AI platform, has been officially released, introducing new features and improvements for enterprise users.

The Pentagon Now Has Its Own Version Of ChatGPT And Grok

The Pentagon has created proprietary AI models similar to ChatGPT and Grok, marking a significant step in military AI capabilities amid rising interest.

Muse Spark 1.3

Meta has announced Muse Spark 1.3, a new update to its AI model, sparking increased interest amid ongoing development and speculation about its capabilities.