AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark shows the lowest-scoring AI still earns 26 points, emphasizing partial progress and trust over perfection. The test highlights key strengths and weaknesses of current models.

A new benchmark developed by Firmulate has demonstrated that even the least effective AI management models score 26 points out of a possible high, revealing critical insights into AI performance under stress. This finding underscores the importance of trust and partial progress in AI-driven management, which are often overlooked in traditional tests.The benchmark involved four frontier AI models managing a simulated small business during a week of crises, customer manipulation attempts, and trust tests. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. Notably, the baseline, representing minimal effort, scored 26 points, illustrating that partial work is valued and recognized in this evaluation. The scoring system emphasizes that trust breaches negate high performance, regardless of competence. The models that successfully identified crucial internal documents and refused social engineering attacks performed better, with some closing high-value deals. Interestingly, even thorough models with deep rule sets, like Opus 4.8, failed to follow through consistently, highlighting that discipline and follow-up are distinct skills. The benchmark’s design aims to measure real-world readiness of AI managers, focusing on trustworthiness, follow-through, and the ability to handle pressure, rather than just conversational ability.
At a glance
reportWhen: published July 2026
The developmentA new benchmark designed to evaluate AI managers’ ability to handle a company’s worst week has revealed surprising results, with the lowest scorer earning 26 points.
Why The Worst AI Manager Still Gets 26 Points
AI Benchmark Report · July 2026 · Firmulate

Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros

Four frontier AI models managed a simulated small business through its worst week — crises, manipulation attempts, and trust tests. Even the floor score of 26 reveals a deliberate design choice: partial progress counts, but trust breaches erase everything.

95 / 100
Top Score — gpt-5.6-sol
73
Lowest Model — Opus 4.8
26
Baseline — Minimum Viable Management
4
Frontier Models
7 days
Simulated Crisis Week
<100
Score Cap — Anti-Inflation
1
Trust Breach = Score Cap
01 · The Scoreboard

Performance Under the Worst Week

The bar chart below maps each model’s final score. The hatched bar represents the baseline: minimal effort — triaging crises and reading documents — is still worth 26 points by design.

GPT-5.6-SOL
95
FRONTIER MODEL 2
86
FRONTIER MODEL 3
81
OPUS 4.8
73
BASELINE (MIN. EFFORT)
26
02 · Evaluation Framework

What the Benchmark Actually Measures

Unlike traditional tests of conversational ability, the Firmulate benchmark scores operational readiness: trustworthiness, follow-through, and resilience under pressure.

Pillar 01 · Trust

Trustworthiness Under Attack

Models faced impersonation and social engineering attempts. A single trust breach caps the maximum score regardless of technical competence — no amount of good work outweighs it.

Pillar 02 · Follow-Through

Discipline vs. Knowledge

Even thorough models with deep rule sets, like Opus 4.8, failed to follow through consistently. Knowing the rules and executing them are distinct skills.

Pillar 03 · Awareness

Internal Awareness

Models that identified and referenced crucial internal documents closed more high-value deals — internal context awareness directly drives performance.

03 · Model Comparison

Strengths and Failure Modes

Capability GPT-5.6-Sol (95) Frontier Group (73–86) Opus 4.8 (73)
Refused social engineering✓ Consistent~ Mostly✓ Consistent
Read internal documentation✓ Deep~ Partial✓ Deep
Closed high-value deals✓ Multiple~ Some✗ Rare
Followed through consistently✓ Strong~ Mixed✗ Inconsistent
Resilience under pressure✓ High~ Moderate~ Moderate
04 · From the Researchers

The Philosophy Behind the Floor

“A manager who does something useful is not the same as one who does nothing, and partial progress counts. The floor isn’t charity; it’s an accounting of what minimum viable management is actually worth.”

— Anonymous Researcher

“No amount of good work outweighs a breach of trust. Even a brilliant performance for six days can be invalidated by a single trust breach.”

— Anonymous Researcher
05 · Roadmap

Next Steps for Benchmark Development

The benchmark is set to evolve into a standard tool for evaluating AI management readiness. Here’s the trajectory:

1
More Diverse Scenarios
2
Longer Management Periods
3
Real-World Deployments
4
Refined Trust Scoring
5
Late-2026 Standard
06 · Key Questions

Answering the Open Debates

Why does the lowest score not equal zero?

The benchmark recognizes partial work as valuable. Even minimal effort — triaging crises and reading documents — earns points, with 26 representing minimum viable management.

What is the significance of trust breaches?

Trust breaches such as impersonation or manipulation cap the maximum score regardless of other performance. Integrity outweighs technical competence.

Do models that read internal documentation perform better?

Yes. Models that successfully read and reference internal files close more deals and score higher, demonstrating the importance of internal awareness.

Will future models score higher?

Uncertain. Models may improve at trust and follow-through, but the benchmark design aims to keep integrity and partial progress at the center of evaluation.

Impact of Partial Progress and Trust in AI Management

This benchmark shifts the focus from AI conversational skills to management qualities critical in real business environments, such as trustworthiness, follow-through, and resilience under pressure. It reveals that even the weakest AI managers contribute meaningful value, but breaches of trust can negate their efforts. For enterprises deploying AI in management roles, these findings highlight the importance of selecting models that prioritize integrity and follow-up, not just technical proficiency. The results also challenge the assumption that perfection is necessary for effective AI management, emphasizing that partial but trustworthy work can be highly valuable. As AI begins to take on more operational roles, understanding these dynamics will influence deployment strategies and risk management, making this benchmark a vital reference for future AI integration.

AI management simulation tools

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Benchmarks and Stress Testing

Traditional AI benchmarks focus primarily on language proficiency, generation quality, or specific task accuracy. Few tests evaluate how well AI models manage ongoing business processes, especially under adverse conditions. The Firmulate benchmark was created to simulate a realistic management scenario, including crises, manipulation attempts, and trust challenges, to assess AI models’ operational robustness. The concept originated from the recognition that real-world AI management involves more than just conversation—it requires follow-through, trustworthiness, and resilience. The July 2026 results are the first comprehensive public test of AI management in a simulated ‘worst week’ scenario, revealing both strengths and weaknesses that are critical for enterprise adoption.

“A manager who does something useful is not the same as one who does nothing, and partial progress counts. The floor isn’t charity; it’s an accounting of what minimum viable management is actually worth.”

— an anonymous researcher

Unanswered Questions About Model Performance and Scoring

It remains unclear how the scoring system will evolve as models improve and more complex scenarios are tested. The benchmark’s design intentionally caps scores at below 100 to prevent inflation, but whether future models will consistently approach higher scores or if trust breaches will remain a decisive factor is still uncertain. Additionally, the real-world applicability of these simulated scenarios needs further validation through live deployments. The specific criteria for trust breaches and how models will be trained to prevent them are still under development, making it uncertain how these results will influence future AI management standards.

Next Steps for Benchmark Development and AI Management Standards

Further testing is planned to include more diverse scenarios, longer management periods, and real-world deployments. Developers will refine the scoring system to better differentiate between partial and full trustworthiness. Industry stakeholders are expected to analyze these results to inform AI deployment strategies, emphasizing integrity and follow-through. The benchmark aims to evolve into a standard tool for evaluating AI management readiness, influencing both model development and enterprise adoption. Researchers and companies will likely collaborate to improve AI’s ability to manage under pressure while maintaining trust, with upcoming updates scheduled for late 2026.

Key Questions

Why does the lowest score in the benchmark not equal zero?

Because the benchmark recognizes partial work as valuable, even minimal effort like triaging crises and reading documents earns points, with 26 representing the minimum viable management effort.

What is the significance of trust breaches in this benchmark?

Trust breaches, such as impersonation or manipulation attempts, caps the maximum score regardless of other performance, emphasizing that integrity is more critical than technical competence.

How do models that read internal documentation perform compared to others?

Models that successfully read and reference internal files tend to close more deals and perform better, demonstrating the importance of internal awareness for effective management.

Will future models be able to score higher than current ones?

It is still uncertain; future models may improve in managing trust and follow-through, but the benchmark’s design aims to maintain a focus on integrity and partial progress.

How can enterprises use this benchmark for their AI deployment?

Enterprises can evaluate their AI models against this benchmark to assess trustworthiness, follow-through, and resilience, informing better deployment decisions and risk management strategies.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A Skill Is a Folder, Not a Prompt: What Anthropic Learned Running Hundreds of Them

Anthropic says reusable Claude Code Skills helped standardize agent work across its engineering organization.

How OBD2 Scanners Interpret Modern Vehicle Data

Knowledge of how OBD2 scanners interpret modern vehicle data reveals crucial insights that can significantly enhance your vehicle’s performance and diagnostics.

Using AI to write better code more slowly

A new approach suggests slowing down AI-assisted coding to improve code quality and bug detection, challenging the fast, low-quality coding narrative.

Maximize Your Notes With These 11 AI-Powered Apps In 2026

A 2026 comparison ranks 11 AI note-taking devices, led by iFLYTEK AINOTE Air 2, while flagging cost and compatibility tradeoffs.