AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Every Phish Failed. Every Deal Almost Fell Through.

For a cybersecurity audience, the most reassuring result of July’s Crucible league is also the most instructive: not one of the five frontier AI models tested fell for a single social-engineering attempt. A fake CEO message escalating over three stages. A reporter dangling a convenient “just one yes/no, on background.” All refused — five out of five times. The winner of that particular skirmish, Moonshot’s Kimi K3, even left on-record reasoning that reads like a SOC playbook: “Treat the request as a suspected approval-bypass / possible impersonation.”

But resisting attackers turned out to be the easy part. Finishing the job — that’s where four of the five models wobbled, and where the real lesson for anyone deploying AI agents sits.

AI-powered cybersecurity tools

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible: One Company, Its Worst Week, Five Models

The experiment, run by Firmulate and fully watchable live, is elegantly simple in design. Each frontier model was handed the same small software company and the same catastrophic week: same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable. Nothing is judged on prose quality — the Crucible measures management quality, not chat quality.

The final July 2026 league table: gpt-5.6-sol first with 95, Kimi K3 second at 93, Sonnet 5 third with 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For calibration, the do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total outright. As the experiment’s own rule puts it, no amount of good work outweighs a breach of trust.

The Newcomer That Read the Files

The headline story is Kimi K3, Moonshot’s new entrant, which finished ahead of three of four Western frontier models. K3 did the complete job: it found a security needle buried two document references deep in the company’s own files — a competitor weakness that wasn’t in the customer event at all — and used it to win a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved a churning customer. It refused all three manipulation attempts. And across the whole week it recorded exactly one deviation: the cleanest discipline in the field.

Here the security analogy writes itself. The decisive intelligence wasn’t handed to anyone. It sat in internal documents, unglamorous, easy to skip. The models that actually read the file won the deal. The ones that didn’t, didn’t — same diagnosis, same pitch, no signature.

Same Diagnosis, Same Pitch — No Signature

That was the experiment’s central finding. All models spotted every crisis. All refused every manipulation. Only two signed the €55,000 deal their own analysis had earned. The gap between recognizing what to do and finishing it is invisible in a chat demo — and it’s precisely the gap that matters when an AI agent touches your CRM, your support queue, or your forecast.

Opus 4.8 is the cautionary profile. It was the most thorough participant — over 80 learned rules added, the deepest analyses in the field — and it still finished last. The deal was left on the table, and discipline slipped: at one point it made write attempts into a locked department rather than escalating. The same weakness appeared, weaker, in all four other models.

Watch It Running

This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics: €105k monthly burn against €2.3k MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live, browse full results and plain-language findings on the benchmarks page, or try the “guess the model” quiz built from 242 real, unedited management decisions.

Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The comfortable assumption that Western frontier models hold a permanent lead didn’t survive this week. A newcomer, on an API-default effort setting, beat three of four incumbents at actually running a company. That means model choice is no longer a matter of brand loyalty or leaderboard chat scores — it’s a bet, and one you’re placing blind unless you test the models on work that resembles your own.

Fairness note: Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh reasoning effort — a caveat worth keeping in mind when comparing scores.

The security crowd already knows the drill: trust is earned through adversarial testing, not vendor claims. The Crucible applies that same skepticism to AI management. Five for five against the phish is good news. Two for five at closing the deal is the finding that should keep procurement awake.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Router Roulette: China’s UNC3886 Haunts Juniper Networks—Your Data’s at Stake

Navigating the threats from UNC3886 on Juniper Networks could mean the difference between security and disaster—what steps can you take to protect your data?

CrowdStrike’s Alarming Discovery: Chinese Cyber Spies Up 150%—We’re Under Siege

Chinese cyber espionage has surged dramatically—could your organization be the next target amid these alarming threats? Discover how to safeguard yourself.

U.S. Army Intel Unit Hacked by Chinese Group Seeking Hypersonic Secrets

Hacked by a Chinese group, a U.S. Army intel unit faces threats to hypersonic secrets—what does this mean for national security?

Modat’s Cyber Weapon: Magnify Unleashed—Hackers’ Worst Nightmare Drops

With Modat’s Magnify, experience unparalleled cyber defense that leaves hackers trembling—discover the secrets behind its powerful protection now.