
Every security team knows the drill: a fake CEO message escalates over three stages, then a “reporter” calls asking for just one yes-or-no answer, on background. It’s textbook social engineering, the kind that empties bank accounts every quarter. When four frontier AI models each spent a week running the same small software company through its worst days, all of them — every single one — refused the bait.
That’s the good news from the final July 2026 standings of the Crucible League, a live benchmark that runs AI models as complete companies rather than chat windows. The bad news is more subtle, and it’s the part that should worry anyone deploying agents near a CRM, a support queue, or a forecast: the models that failed didn’t fail at staying honest. They failed at doing their homework. A decisive competitive fact was sitting in the company’s own files, two document references deep — and only the agents that actually read it closed the deal.
Same crisis week, same temptations
The setup, run by Firmulate, is refreshingly blunt. Each frontier model got the identical job: run a 13-employee synthetic software company through a brutal week — the same customers, the same crises, the same chances to cheat. Real money mechanics apply: the company burns €105,000 a month against just €2,300 in monthly recurring revenue, and every workday is versioned and auditable, so no model can quietly edit its past. The public can watch it happen at firmulate.com/live, complete with an open cash countdown.
The final league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. For calibration, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total outright. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”
CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The manipulation test: five for five
For a privacy-and-espionage audience, the social-engineering results are the headline. The fake CEO messages escalated across three stages; a fake reporter tried the classic “just one yes/no, on background” squeeze. Five of five models refused. Kimi K3 left its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That’s a meaningful result. AI agents are increasingly the entity answering the phone, the email, and the chat widget — which makes them the target of exactly the impersonation plays humans fall for. On this evidence, frontier models are harder targets than the average finance clerk. The failure mode isn’t gullibility.
The buried fact: intelligence work, not charm
Here’s where the story turns into an espionage tale told in reverse. The €55,000 deal at the center of the week hinged on a competitor weakness — and that weakness wasn’t in the customer conversation at all. It sat in the company’s own files, two document references deep. An agent had to follow one document to another to find it.
The models split cleanly on this. All of them diagnosed the crisis correctly. All of them made the pitch. Only two actually signed the deal their own analysis had earned — the experiment’s blunt summary: “Same diagnosis, same pitch — no signature.” The winners closed at full price, worth +€4,583 in monthly recurring revenue. The losers left it on the table automatically.
In intelligence terms, the winners did collection before analysis. The losers skipped to the assessment. And nothing about a chat demo would ever reveal the difference — which is precisely the point of running agents through a full business week rather than a prompt showcase.
The thoroughness paradox
The most instructive profile belongs to Opus 4.8: the most thorough participant in the field, contributing more than 80 self-learned playbook rules and producing the deepest analyses — and finishing dead last. The close was never made, and discipline slipped in a way that will sound familiar to anyone who has watched an overzealous insider go around process: it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four competitors. Diligence and compliance are not the same property, and this benchmark scores both.
One fairness footnote worth noting: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still took second place with the cleanest discipline of the field.
Why it’s watchable, and why it’s yours to test
Firmulate keeps the whole thing in public view. The live company now carries more than 680 self-learned playbook rules, the site rebuilds itself twice a day, and 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. Full results and plain-language findings are published at firmulate.com/benchmarks.html.
For enterprises, there’s a pilot with a security-friendly design: run the same wargame against a read-only export of your own business. Nothing ever writes back to real systems — the agent gets the files, the pressure, and the temptations, and your infrastructure stays untouched. Details are at firmulate.com/pilot.html.

The security lesson from the Crucible League cuts against the usual fear narrative. Frontier AI agents resisted every impersonation and manipulation attempt thrown at them — three-stage fake-CEO escalations, reporter traps, the works. What separated the winners from the losers was a humbler skill: whether the agent read the files in front of it before answering. That’s a measurable, purchase-deciding property, and it cost two of the field’s models a €55,000 deal that their own analysis had already earned.
So before you trust an agent with your inbox, your queue, or your pipeline, ask a different question than “is it safe?” Ask: does it do its homework? Because the buried fact in your own documents is exactly where a competitor’s agent — or your own — will either find the win or leave it for someone who looked one reference deeper.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html