AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A fake CEO message arrives. A reporter asks for “just one yes/no, on background.” The obvious test is whether an AI gives away secrets. But a harder question follows: after it resists the trick, can it still do the job it was trusted to do?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That question sits at the center of Firmulate, a live experiment in which AI models run a small software company through a week of crises. The test links familiar security concerns—impersonation, manipulation and access discipline—to a less visible risk: an AI can recognize trouble and still fail to act well.

A security test with business consequences

In the final Crucible League, published in July 2026, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. In a three-stage sequence of fake CEO messages, followed by the reporter trick, all five models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”

That is a strong first line of defense. Yet the experiment’s most consequential gap appeared after the threats had been recognized. Only two models signed a €55,000 deal that their own analysis had earned. Firmulate summed up the result: “Same diagnosis, same pitch — no signature.” The finding is a reminder that security behavior cannot be judged only by whether a model refuses a suspicious request. Its handling of legitimate work matters too.

The clue was already inside the company

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result turns a familiar information-security concern into a practical management question: can an AI find and use relevant material in the business records it is allowed to inspect, while keeping its actions within bounds?

The league rankings make the performance differences tangible: gpt-5.6-sol placed first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate’s rule is uncompromising: partial progress counts, but one breach of trust caps the total—“no amount of good work outweighs a breach of trust.”

Opus 4.8 offers a particularly sharp illustration of why a single security score is not the whole story. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but finished last. It left the deal unsigned and discipline slipped when it attempted writes in a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. Thoroughness and caution are useful; neither guarantees sound execution.

There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The published ranking should be read with that difference in mind.

A company you can watch

Firmulate describes its live company as 13 synthetic employees operating with real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and workdays that are versioned. The live experiment is watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites visitors to guess which model made them.

For security and privacy teams, the appeal is the ability to observe decisions under pressure and examine where playbooks hold—or fail. The experiment moves the conversation beyond polished chat demonstrations toward how an AI handles competing demands: protect trust, follow its own analysis, find the relevant evidence and escalate when it lacks authority.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to testing your own business

Firmulate’s enterprise pilot applies the wargame to a read-only export of a company’s own business. Teams can explore crisis scenarios and review a board report with model rankings and weak points in their playbooks. The pilot does not write back to real systems.

To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

GCHQ’s AI Triumph: Foils Cyber Espionage on UK Defense Systems

Unveiling GCHQ’s groundbreaking AI strategies reveals how they thwart cyber espionage, but the full extent of their innovations remains to be explored.

Counterintelligence Chief: A.I. Deepfakes Could Fuel Next Wave of Terror Propaganda

Grappling with the rise of AI deepfakes, experts warn they could ignite a new wave of terror propaganda, but how can we effectively counter this threat?

Chinese APT Group Enhances European Cyber Ops With Advanced Malware

Get ready to uncover how Chinese APT groups are revolutionizing European cyber operations with advanced malware, but what does this mean for your security?

MI6’s AI Find: Russian Oligarchs Fund UK Espionage via Property

Learn how Russian oligarchs are secretly funding UK espionage through property investments, raising alarming questions about national security and government action. What lies beneath?