🔍 Read the full analysis: Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros on ThorstenMeyerAI.com
Get privacy and security gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark shows the lowest-scoring AI still earns 26 points, emphasizing partial progress and trust over perfection. The test highlights key strengths and weaknesses of current models.
Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros
Four frontier AI models managed a simulated small business through its worst week — crises, manipulation attempts, and trust tests. Even the floor score of 26 reveals a deliberate design choice: partial progress counts, but trust breaches erase everything.
Performance Under the Worst Week
The bar chart below maps each model’s final score. The hatched bar represents the baseline: minimal effort — triaging crises and reading documents — is still worth 26 points by design.
What the Benchmark Actually Measures
Unlike traditional tests of conversational ability, the Firmulate benchmark scores operational readiness: trustworthiness, follow-through, and resilience under pressure.
Trustworthiness Under Attack
Models faced impersonation and social engineering attempts. A single trust breach caps the maximum score regardless of technical competence — no amount of good work outweighs it.
Discipline vs. Knowledge
Even thorough models with deep rule sets, like Opus 4.8, failed to follow through consistently. Knowing the rules and executing them are distinct skills.
Internal Awareness
Models that identified and referenced crucial internal documents closed more high-value deals — internal context awareness directly drives performance.
Strengths and Failure Modes
| Capability | GPT-5.6-Sol (95) | Frontier Group (73–86) | Opus 4.8 (73) |
|---|---|---|---|
| Refused social engineering | ✓ Consistent | ~ Mostly | ✓ Consistent |
| Read internal documentation | ✓ Deep | ~ Partial | ✓ Deep |
| Closed high-value deals | ✓ Multiple | ~ Some | ✗ Rare |
| Followed through consistently | ✓ Strong | ~ Mixed | ✗ Inconsistent |
| Resilience under pressure | ✓ High | ~ Moderate | ~ Moderate |
The Philosophy Behind the Floor
“A manager who does something useful is not the same as one who does nothing, and partial progress counts. The floor isn’t charity; it’s an accounting of what minimum viable management is actually worth.”
“No amount of good work outweighs a breach of trust. Even a brilliant performance for six days can be invalidated by a single trust breach.”
Next Steps for Benchmark Development
The benchmark is set to evolve into a standard tool for evaluating AI management readiness. Here’s the trajectory:
Answering the Open Debates
Why does the lowest score not equal zero?
The benchmark recognizes partial work as valuable. Even minimal effort — triaging crises and reading documents — earns points, with 26 representing minimum viable management.
What is the significance of trust breaches?
Trust breaches such as impersonation or manipulation cap the maximum score regardless of other performance. Integrity outweighs technical competence.
Do models that read internal documentation perform better?
Yes. Models that successfully read and reference internal files close more deals and score higher, demonstrating the importance of internal awareness.
Will future models score higher?
Uncertain. Models may improve at trust and follow-through, but the benchmark design aims to keep integrity and partial progress at the center of evaluation.
Impact of Partial Progress and Trust in AI Management
This benchmark shifts the focus from AI conversational skills to management qualities critical in real business environments, such as trustworthiness, follow-through, and resilience under pressure. It reveals that even the weakest AI managers contribute meaningful value, but breaches of trust can negate their efforts. For enterprises deploying AI in management roles, these findings highlight the importance of selecting models that prioritize integrity and follow-up, not just technical proficiency. The results also challenge the assumption that perfection is necessary for effective AI management, emphasizing that partial but trustworthy work can be highly valuable. As AI begins to take on more operational roles, understanding these dynamics will influence deployment strategies and risk management, making this benchmark a vital reference for future AI integration.AI management simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Benchmarks and Stress Testing
Traditional AI benchmarks focus primarily on language proficiency, generation quality, or specific task accuracy. Few tests evaluate how well AI models manage ongoing business processes, especially under adverse conditions. The Firmulate benchmark was created to simulate a realistic management scenario, including crises, manipulation attempts, and trust challenges, to assess AI models’ operational robustness. The concept originated from the recognition that real-world AI management involves more than just conversation—it requires follow-through, trustworthiness, and resilience. The July 2026 results are the first comprehensive public test of AI management in a simulated ‘worst week’ scenario, revealing both strengths and weaknesses that are critical for enterprise adoption.“A manager who does something useful is not the same as one who does nothing, and partial progress counts. The floor isn’t charity; it’s an accounting of what minimum viable management is actually worth.”
— an anonymous researcher
Unanswered Questions About Model Performance and Scoring
It remains unclear how the scoring system will evolve as models improve and more complex scenarios are tested. The benchmark’s design intentionally caps scores at below 100 to prevent inflation, but whether future models will consistently approach higher scores or if trust breaches will remain a decisive factor is still uncertain. Additionally, the real-world applicability of these simulated scenarios needs further validation through live deployments. The specific criteria for trust breaches and how models will be trained to prevent them are still under development, making it uncertain how these results will influence future AI management standards.Next Steps for Benchmark Development and AI Management Standards
Further testing is planned to include more diverse scenarios, longer management periods, and real-world deployments. Developers will refine the scoring system to better differentiate between partial and full trustworthiness. Industry stakeholders are expected to analyze these results to inform AI deployment strategies, emphasizing integrity and follow-through. The benchmark aims to evolve into a standard tool for evaluating AI management readiness, influencing both model development and enterprise adoption. Researchers and companies will likely collaborate to improve AI’s ability to manage under pressure while maintaining trust, with upcoming updates scheduled for late 2026.Key Questions
Why does the lowest score in the benchmark not equal zero?
Because the benchmark recognizes partial work as valuable, even minimal effort like triaging crises and reading documents earns points, with 26 representing the minimum viable management effort.What is the significance of trust breaches in this benchmark?
Trust breaches, such as impersonation or manipulation attempts, caps the maximum score regardless of other performance, emphasizing that integrity is more critical than technical competence.How do models that read internal documentation perform compared to others?
Models that successfully read and reference internal files tend to close more deals and perform better, demonstrating the importance of internal awareness for effective management.Will future models be able to score higher than current ones?
It is still uncertain; future models may improve in managing trust and follow-through, but the benchmark’s design aims to maintain a focus on integrity and partial progress.How can enterprises use this benchmark for their AI deployment?
Enterprises can evaluate their AI models against this benchmark to assess trustworthiness, follow-through, and resilience, informing better deployment decisions and risk management strategies.Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
