
Get privacy and security gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Every Phish Failed. Every Deal Almost Fell Through.
For a cybersecurity audience, the most reassuring result of July’s Crucible league is also the most instructive: not one of the five frontier AI models tested fell for a single social-engineering attempt. A fake CEO message escalating over three stages. A reporter dangling a convenient “just one yes/no, on background.” All refused — five out of five times. The winner of that particular skirmish, Moonshot’s Kimi K3, even left on-record reasoning that reads like a SOC playbook: “Treat the request as a suspected approval-bypass / possible impersonation.”
But resisting attackers turned out to be the easy part. Finishing the job — that’s where four of the five models wobbled, and where the real lesson for anyone deploying AI agents sits.
AI-powered cybersecurity tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Crucible: One Company, Its Worst Week, Five Models
The experiment, run by Firmulate and fully watchable live, is elegantly simple in design. Each frontier model was handed the same small software company and the same catastrophic week: same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable. Nothing is judged on prose quality — the Crucible measures management quality, not chat quality.
The final July 2026 league table: gpt-5.6-sol first with 95, Kimi K3 second at 93, Sonnet 5 third with 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For calibration, the do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total outright. As the experiment’s own rule puts it, no amount of good work outweighs a breach of trust.
The Newcomer That Read the Files
The headline story is Kimi K3, Moonshot’s new entrant, which finished ahead of three of four Western frontier models. K3 did the complete job: it found a security needle buried two document references deep in the company’s own files — a competitor weakness that wasn’t in the customer event at all — and used it to win a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved a churning customer. It refused all three manipulation attempts. And across the whole week it recorded exactly one deviation: the cleanest discipline in the field.
Here the security analogy writes itself. The decisive intelligence wasn’t handed to anyone. It sat in internal documents, unglamorous, easy to skip. The models that actually read the file won the deal. The ones that didn’t, didn’t — same diagnosis, same pitch, no signature.
Same Diagnosis, Same Pitch — No Signature
That was the experiment’s central finding. All models spotted every crisis. All refused every manipulation. Only two signed the €55,000 deal their own analysis had earned. The gap between recognizing what to do and finishing it is invisible in a chat demo — and it’s precisely the gap that matters when an AI agent touches your CRM, your support queue, or your forecast.
Opus 4.8 is the cautionary profile. It was the most thorough participant — over 80 learned rules added, the deepest analyses in the field — and it still finished last. The deal was left on the table, and discipline slipped: at one point it made write attempts into a locked department rather than escalating. The same weakness appeared, weaker, in all four other models.
Watch It Running
This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics: €105k monthly burn against €2.3k MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live, browse full results and plain-language findings on the benchmarks page, or try the “guess the model” quiz built from 242 real, unedited management decisions.
Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The League Is Open
The comfortable assumption that Western frontier models hold a permanent lead didn’t survive this week. A newcomer, on an API-default effort setting, beat three of four incumbents at actually running a company. That means model choice is no longer a matter of brand loyalty or leaderboard chat scores — it’s a bet, and one you’re placing blind unless you test the models on work that resembles your own.
Fairness note: Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh reasoning effort — a caveat worth keeping in mind when comparing scores.
The security crowd already knows the drill: trust is earned through adversarial testing, not vendor claims. The Crucible applies that same skepticism to AI management. Five for five against the phish is good news. Two for five at closing the deal is the finding that should keep procurement awake.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
