
The Phish That Didn’t Land — and the Failure That Still Happened
If you work in security, you’ve seen this movie: an urgent message from the CEO, escalating pressure, then a journalist offering a quick “just one yes/no, on background” to make it all go away. It’s textbook social engineering. When researchers at Firmulate ran four frontier AI models through exactly that gauntlet — each one handed the same small software company for its worst week — the manipulation attempts failed completely. Five out of five attempts, refused across the board. One model, Kimi K3, even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
So the bots passed the phishing test. Case closed? Not quite. Because in the same week, three of those four models still managed to fail the business — in a way no chat benchmark would ever catch.
AI-powered phishing detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment
Firmulate’s Crucible League gave each frontier model an identical job: run the same small software company through a week of crises — a churn wave, a price increase, a downround, a PR blowup — with the same customers, the same temptations to cheat, and every decision versioned and auditable. Only the model changed. The final July 2026 standings: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total outright. As the rules put it: no amount of good work outweighs a breach of trust.
Same Diagnosis, Same Pitch — No Signature
Here’s the finding that matters. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The companies got the right answer and still didn’t get the money.
And the buried fact behind the winners: the decisive competitor weakness wasn’t in the customer’s event at all. It sat two document references deep in the company’s own files. The models that actually read their own documents won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
That’s a security-adjacent lesson in itself: the most valuable intelligence was already inside the perimeter, in files nobody read.
The Thoroughness Trap
The most striking profile belongs to Opus 4.8: the most thorough participant in the field, generating 80 additional learned rules and the deepest analyses — and still finishing last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. From a governance standpoint, that’s the uncomfortable part — the same weakness appeared, weaker, in all four models. Diligence and discipline are not the same axis, and only one of them shows up in a demo.
One fairness note Firmulate discloses openly: K3 ran at API-default effort while the others ran at xhigh — and still came second.
Not a Slide Deck
This isn’t a one-off paper. The company is live software with 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it lose money in real time at firmulate.com. There’s also a quiz built from 242 real, unedited management decisions where you guess which model made which call.
For enterprises, Firmulate offers a pilot: run the same wargame against a read-only export of your own business. Nothing ever writes back to real systems — a constraint any security team will appreciate.

Management Quality, Not Chat Quality
The security industry spent a decade teaching humans not to click. The models have, reassuringly, learned that lesson. What they haven’t been measured on — until now — is whether they finish what they start, read the files they already have, and stay disciplined when nobody’s escalating on their behalf.
Leaderboards and chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences across days, or what happens when the deal is on the table and the document trail is two references deep. If an AI agent will touch your CRM, your support queue, or your forecast, “does it write well” is the wrong question. The right one: does it manage? As Firmulate puts it — management quality, not chat quality. The gap between the two is exactly where three of four frontier models just fell in.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html