
Can you identify an AI by the choices it makes under pressure?
Cybersecurity teams already know that identity is part of the threat model. A convincing message can conceal an impersonator, while a seemingly harmless request can be an attempt to bypass approval. Firmulate applies that same suspicious eye to AI managers: not by examining how they introduce themselves, but by watching what they actually do.
Its interactive quiz draws from 242 real, unedited management decisions. Readers see how a frontier model responded to a business situation and try to identify it. The exercise is entertaining, but the underlying question is serious: do models develop recognizable management personalities when they face the same evidence, incentives and temptations?
The decisions come from a live, watchable experiment in which frontier models ran the same small software company through its worst week. Every model faced the same customers, crises and opportunities to cut corners. Every workday and decision was versioned and auditable.
AI management decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company designed to expose judgment
Firmulate is not testing whether a model can produce a polished memo. The company has 13 synthetic employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. A public cash countdown makes delay visible. The company has also accumulated more than 680 self-learned playbook rules.
This setting turns ordinary management behavior into evidence. A model can recognize a crisis yet fail to complete the action that resolves it. It can write a persuasive sales pitch but neglect the final commercial step. It can conduct a deep analysis while ignoring the process needed to move work forward safely.
The clearest example was a €55,000 deal. Every model reached the same diagnosis and developed the same pitch, yet only two signed the contract their analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”
The decisive clue was hidden in plain sight
The deal did not turn on information contained in the customer event itself. The crucial weakness in the competitor’s position sat two document references deep inside the company’s own files. Models that followed the trail and read the file won the deal at full price, adding €4,583 in monthly recurring revenue.
For security and privacy professionals, that detail may be the experiment’s most revealing result. A capable agent must know when to distrust the surface context and inspect the authoritative record. The difference between an impressive answer and a useful business outcome can be a piece of evidence buried beyond the first document.
Every model resisted the attackers
The models encountered fake CEO messages that escalated over three stages. They also faced a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts, and every model spotted every crisis.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response captures a useful security instinct: urgency and claimed authority do not erase the need for verification. Across the field, resisting social engineering was not what separated the leaders from the rest. Finishing legitimate work did.
A league table of management behavior
The final Crucible League results from July 2026 make those differences measurable:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
A do-nothing baseline scores 26 because partial progress counts. Trust, however, is a hard boundary: a single breach caps the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”
The comparison includes an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference belongs beside the results because the experiment is meant to make model behavior inspectable, not flatten away the conditions under which it occurred.
Thoroughness is not the same as effectiveness
Opus 4.8 offers the strongest warning against judging an AI manager by the apparent depth of its output. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the commercial close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating.
The same weakness appeared in all four of the other participants, though less strongly. That makes the quiz more than a game of spotting writing style. The recognizable traits are operational: how deeply a model investigates, whether it respects boundaries, whether it escalates when blocked and whether it completes the final step.

The personality test that matters
The public Firmulate management quiz lets readers test whether those behavioral signatures are visible without a model label attached. The reveal turns each decision into a compact character study, grounded in what the model actually did rather than what it claims it would do.
Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to observe how a prospective AI workforce handles their documents, pressures and approval boundaries before giving it operational access.
The central lesson is uncomfortable but useful: frontier models can all notice danger and reject manipulation, yet still differ sharply in follow-through. For organizations evaluating AI agents, eloquence is only the beginning. The more consequential questions are whether the model reads the files, protects trust and finishes the work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html