🔍 Read the full analysis: Why The Most Flawed AI Managers Still Achieve A 26-Point Score on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A recent AI management benchmark shows even the most flawed AI managers score at least 26 points out of 100. The scoring reflects minimal viable management and trust considerations, not just performance.
The recent results from the Firmulate benchmark league reveal that the lowest-scoring AI managers still receive a score of 26 out of 100, even when they perform minimal or no work. This finding underscores the benchmark’s focus on trust and fundamental management functions, rather than pure performance metrics. For a detailed analysis, see the original analysis. The results highlight that even flawed AI managers contribute some value, and that trust and integrity are decisive factors in scoring, making this a significant update for enterprise AI deployment. For more insights, see the detailed coverage in the original analysis.
The benchmark tested four frontier AI models managing a small software company through a simulated worst week, with identical conditions, crises, and temptations to cut corners. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. Notably, the baseline “do-nothing” model scored 26, emphasizing that even minimal effort is recognized as partial management. This baseline score reflects the minimal work necessary to maintain basic operations, such as triaging customer issues or reading inboxes, which the benchmark considers the lowest viable management contribution.
The scoring system is designed to prevent grade inflation: a perfect 100 is considered suspicious, as it would suggest unmeasured or overly idealized behavior. This approach aligns with broader AI management benchmarks that emphasize trust and integrity. Instead, the 26-point baseline demonstrates that partial, honest effort is valued, but trust breaches—such as failing to escalate or follow through—immediately limit the maximum possible score. The benchmark’s principle is clear: “no amount of good work outweighs a breach of trust,” making integrity a critical component of AI management performance.
Among the models, those that read documentation thoroughly and avoided manipulation attempts scored higher, with the top models successfully closing high-value deals and resisting social engineering attacks. For example, two models identified critical documents buried in the company files, enabling them to win a €55,000 deal, while others failed to do so. Despite their technical sophistication, models that lacked discipline or follow-through, even with extensive rule sets, finished lower, illustrating that thoroughness alone does not guarantee success.
Why The Most Flawed AI Managers Still Achieve A 26-Point Score
Four frontier AI models managed a small software company through a simulated worst week — identical conditions, crises, and temptations to cut corners. Even the “do-nothing” baseline earns 26 of 100 points, revealing a scoring philosophy built on trust, not just performance.
Every Manager Beats The Floor — None Reach Perfection
The striped baseline represents minimal viable management: triaging customer issues, reading inboxes, and maintaining basic operations. Every model clears it — but only disciplined, thorough models climb far above it.
Trust Sets The Ceiling, Effort Sets The Floor
The scale below shows how the benchmark distributes value: partial, honest effort is always recognized — but a single breach of trust immediately caps the maximum achievable score, no matter how impressive the work.
Partial Credit Is Real Credit
Basic tasks like triaging issues and reading inboxes count as the lowest viable management contribution — the 26-point floor rewards showing up honestly.
A Perfect 100 Is Suspicious
The scoring system treats flawless scores as a red flag for unmeasured or overly idealized behavior — grade inflation is designed out by default.
Trust Breaches Cap Scores
Failing to escalate or follow through immediately limits the maximum possible score. Integrity failures cannot be offset by technical brilliance.
Thoroughness And Resistance — Not Just Capability
| Behavior Tested | Top Performers | Lower Scorers | Outcome Impact |
|---|---|---|---|
| Reading buried documentation | ✓ Found critical files | ✗ Missed key documents | €55,000 deal won or lost |
| Resisting social engineering | ✓ Rejected manipulation | ~ Mixed results | Trust score preserved |
| Closing high-value deals | ✓ Followed through | ~ Inconsistent follow-up | Direct score gains |
| Rule-set adherence under pressure | ✓ Disciplined execution | ✗ Lapsed despite extensive rules | Thoroughness alone insufficient |
| Escalating crises | ✓ Prompt escalation | ✗ Failed to escalate | Score cap triggered |
The Benchmark’s Non-Negotiable Rule
“No amount of good work outweighs a breach of trust.”
Firmulate Benchmark League — Scoring DoctrineSimulate Worst Week
Identical crises, temptations, and adversarial conditions for every model.
Audit Every Decision
A full audit trail of decisions ensures transparency and accountability.
Reward Honest Effort
Partial, verifiable management contributions earn credit from the 26-point floor up.
Cap After Breaches
Any trust breach locks the ceiling — integrity outranks raw performance.
Trustworthy-But-Imperfect Beats Capable-But-Unreliable
For enterprise AI deployment, the key questions shift from raw capability to reliability: Can the agent complete tasks dependably, read relevant documents, and maintain integrity under pressure? The results suggest pragmatic strategies favoring imperfect-but-trustworthy systems.
Pragmatic Over Perfect
Organizations may prefer flawed but trustworthy AI managers over seemingly more capable systems that cannot be relied upon when it matters.
Auditable Decisions
The emphasis on transparent decision-making aligns with growing enterprise demands for accountability as AI enters critical business functions.
Does 26 Generalize?
Whether the baseline reflects real-world minimal management across industries — and how trust breaches are precisely defined — remains under debate.
The 26-Point Threshold, Explained
Why does the baseline score of 26 exist?
It represents the minimum viable management effort — basic tasks like triaging issues and reading inboxes are recognized as partial but essential contributions to management.
What counts as a breach of trust?
Failing to escalate issues, follow through on commitments, or acting dishonestly — any of these immediately caps the maximum score achievable.
Can flawed AI managers still be useful?
Yes. Even the least capable models contribute some management functions, proving partial effort plus trustworthiness carries real-world value.
Will future benchmarks change the scoring?
Likely. Future iterations are expected to explore broader scenarios, more nuanced trust metrics, and real-world validation of the baseline score.
Implications of the 26-Point Baseline for AI Management
This finding matters because it shifts the focus from performance alone to trustworthiness and minimal effective management in AI systems. For enterprise applications, the key questions are whether AI agents can complete tasks reliably, read relevant documents, and maintain integrity under pressure. The benchmark’s scoring system, which awards partial credit for effort but caps scores after trust breaches, encourages deploying AI that is honest and disciplined rather than just technically capable. This approach could influence how organizations evaluate and implement AI management tools, prioritizing trust and consistency over raw performance.
Furthermore, the results suggest that even flawed AI managers contribute value, as long as they avoid breaches of trust. This could lead to more pragmatic deployment strategies, where imperfect but trustworthy AI systems are preferred over seemingly more capable but unreliable ones. The emphasis on auditable, transparent decision-making aligns with increasing enterprise demands for accountability in AI use, especially as these systems become more integrated into critical business functions.
As an affiliate, we earn on qualifying purchases.
Background and Development of the Benchmark
The Firmulate benchmark league was created to evaluate AI managers in realistic business scenarios, focusing on their ability to handle crises, trust attacks, and complex decision-making over a simulated week. Unlike traditional benchmarks that measure language fluency or task-specific accuracy, this test assesses how well AI manages through worst-case conditions, emphasizing trust, follow-through, and integrity. The benchmark’s design includes a full audit trail of decisions, ensuring transparency and accountability.
Since its launch, the league has revealed that even the least capable models perform some management functions, leading to the establishment of a baseline score of 26 points. The results have sparked debate about the true value of partial effort and the importance of trust in AI systems, especially as organizations increasingly rely on AI for critical operations. The benchmark also underscores that perfect scores are unlikely and possibly suspicious, reinforcing the idea that trustworthiness is non-negotiable in AI management.
This development builds on prior research emphasizing that AI systems must be reliable and honest, especially under pressure, and that technical prowess alone is insufficient for real-world deployment. The league’s results serve as a practical measure of how well AI models adhere to these principles in complex, adversarial environments.
enterprise AI trust assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About the 26-Point Threshold
It remains unclear whether the 26-point baseline accurately reflects real-world minimal management efforts across diverse industries or if it is specific to this simulated scenario. Additionally, the precise criteria that define a breach of trust and how these are measured in practice are still evolving. The extent to which partial progress can be reliably scaled or improved upon in operational settings also remains to be seen. Furthermore, the implications of these findings for future AI development and deployment standards are still under discussion.
AI documentation management system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarks and Deployment
The immediate next step is to observe how organizations incorporate these insights into their AI deployment strategies, prioritizing trust and discipline alongside technical capability. Developers and enterprises may refine their evaluation criteria to emphasize trustworthiness, transparency, and follow-through. Future iterations of the benchmark are expected to explore broader scenarios, more nuanced trust metrics, and real-world validation of the baseline score. Additionally, ongoing research will likely investigate how partial effort and trust breaches impact long-term AI integration in critical business functions.
Stakeholders should watch for updates from the benchmark creators, including new scoring models, expanded testing environments, and industry-specific adaptations. Ultimately, the goal is to foster AI systems that are not only capable but also reliable and trustworthy in managing complex, real-world operations.
AI cybersecurity and social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the baseline score of 26 exist in the benchmark?
The score of 26 represents the minimum viable management effort, accounting for basic tasks like triaging issues and reading inboxes, which are recognized as partial but essential contributions to management.
What does a breach of trust mean in this context?
A breach of trust occurs when an AI manager fails to escalate issues, follow through on commitments, or acts dishonestly, which immediately caps the maximum score achievable.
Can flawed AI managers still be useful despite low scores?
Yes, the benchmark shows that even the least capable models can contribute some management functions, emphasizing that partial effort and trustworthiness are valuable in real-world applications.
Will future benchmarks change the scoring system?
It is likely. The current system emphasizes trust and partial effort, but future iterations may refine these criteria or expand scenarios to better reflect operational realities.
How should organizations interpret these results for their AI strategies?
Organizations should prioritize deploying AI systems that demonstrate discipline and trustworthiness, understanding that partial progress is valuable but trust breaches limit overall effectiveness.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
