AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A fake CEO message arrives. A reporter asks for “just one yes/no, on background.” The obvious test is whether an AI gives away secrets. But a harder question follows: after it resists the trick, can it still do the job it was trusted to do?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That question sits at the center of Firmulate, a live experiment in which AI models run a small software company through a week of crises. The test links familiar security concerns—impersonation, manipulation and access discipline—to a less visible risk: an AI can recognize trouble and still fail to act well.

A security test with business consequences

In the final Crucible League, published in July 2026, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. In a three-stage sequence of fake CEO messages, followed by the reporter trick, all five models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”

That is a strong first line of defense. Yet the experiment’s most consequential gap appeared after the threats had been recognized. Only two models signed a €55,000 deal that their own analysis had earned. Firmulate summed up the result: “Same diagnosis, same pitch — no signature.” The finding is a reminder that security behavior cannot be judged only by whether a model refuses a suspicious request. Its handling of legitimate work matters too.

The clue was already inside the company

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result turns a familiar information-security concern into a practical management question: can an AI find and use relevant material in the business records it is allowed to inspect, while keeping its actions within bounds?

The league rankings make the performance differences tangible: gpt-5.6-sol placed first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate’s rule is uncompromising: partial progress counts, but one breach of trust caps the total—“no amount of good work outweighs a breach of trust.”

Opus 4.8 offers a particularly sharp illustration of why a single security score is not the whole story. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but finished last. It left the deal unsigned and discipline slipped when it attempted writes in a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. Thoroughness and caution are useful; neither guarantees sound execution.

There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The published ranking should be read with that difference in mind.

A company you can watch

Firmulate describes its live company as 13 synthetic employees operating with real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and workdays that are versioned. The live experiment is watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites visitors to guess which model made them.

For security and privacy teams, the appeal is the ability to observe decisions under pressure and examine where playbooks hold—or fail. The experiment moves the conversation beyond polished chat demonstrations toward how an AI handles competing demands: protect trust, follow its own analysis, find the relevant evidence and escalate when it lacks authority.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to testing your own business

Firmulate’s enterprise pilot applies the wargame to a read-only export of a company’s own business. Teams can explore crisis scenarios and review a board report with model rankings and weak points in their playbooks. The pilot does not write back to real systems.

To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Intelligence Bulletin: Deepfake Threats Loom Over Upcoming Elections

Understanding deepfake threats is crucial as they pose a significant risk to upcoming elections; uncover how to protect democracy before it’s too late.

China’s AI Frenzy: Hype Machine Accelerates—Global Spy Domination Near?

Uncover the truth behind China’s AI surge and its potential for global surveillance dominance—what implications could this have for the future?

Ukraine’s SBU Exposes Russian Mole in Zelensky’s Inner Circle

Not only has Ukraine’s SBU unveiled a dangerous Russian mole within Zelenskyy’s circle, but the implications for national security are staggering and far-reaching.

Singapore’s ISD Busts Chinese Spy Ring Targeting ASEAN Summit Plans

Fearing for national security, Singapore’s ISD has uncovered a Chinese spy ring targeting ASEAN summit plans, raising alarming questions about regional espionage. What will happen next?