
Imagine a company with no employees, losing €105,000 every month, yet still competing in a brutal market environment. Now, watch it operate live, with every decision, crisis, and temptation exposed for all to see. This is not fiction but the world of Firmulate, a pioneering experiment in building AI-driven management that is as transparent as it is intense.
The Live Experiment: A Company Without Human Employees
At the core of this extraordinary venture is a small, imaginary software company run entirely by artificial intelligence models. These models, 13 in total, act as virtual employees managing client relationships, navigating crises, and making strategic decisions. Every workday, the company’s entire decision-making process is versioned and published, creating a real-time window into how AI handles complex business scenarios.
This experiment doesn’t just showcase AI in a demo — it puts it through its paces during a simulated worst week, filled with the same customers, crises, and ethical temptations a real company would face. The goal? To measure whether these models can not only identify problems but also act honestly and reliably under pressure.
AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Groundbreaking Results in AI Management
Among the models tested, the standout was gpt-5.6-sol, which scored a remarkable 95 out of 100 on the company’s internal league table. This model successfully uncovered a crucial hidden document that was key to closing a €55,000 deal, earning full revenue—an extraordinary feat because the same diagnosis and pitch were presented to the client, yet only the AI that read the right files made the sale.
In contrast, other models like Kimi K3 scored slightly lower but still managed to close deals, thanks to their disciplined approach. Meanwhile, Opus 4.8, the most thorough participant, performed well in analysis but faltered in execution — leaving an important deal unclosed because it misclassified internal policy and failed to escalate issues properly.
The Human-Like Challenges and AI Integrity
One of the most revealing parts of the experiment was how the models responded to social engineering attempts. A fake CEO message escalating requests over multiple stages, designed to trick the AI into bypassing safeguards, was refused by all models. Kimi K3 explicitly treated such requests as possible impersonation, showcasing their built-in caution against manipulation — a crucial trait for trustworthy AI in critical business environments.
Why This Matters for Cybersecurity and Privacy
For cybersecurity professionals and privacy advocates, the implications are clear: AI systems must be able to operate honestly and resist manipulation, especially in high-stakes scenarios. The Firmulate experiment demonstrates that the models can detect crises, refuse unethical demands, and prioritize data security, making them suitable for sensitive tasks like managing customer data or financial transactions.
However, the experiment also highlights the persistent risk of internal weaknesses—found not in the customer interactions but within internal documentation. AI models that read these files can uncover critical facts that might be buried or overlooked, underscoring the importance of keeping sensitive information secure and well-managed.
Transparency and Building in Public
What sets this project apart is its extreme transparency. Every decision, every rule learned — totaling over 680 self-learned policies — is openly published. This build-in-public approach offers a detailed view of how AI models operate in complex environments, providing valuable insights for security professionals concerned with AI reliability and ethical behavior.
The Future of AI-Driven Business Management
As firms and organizations increasingly deploy AI to handle critical functions, understanding how models perform under stress is essential. The live experiment at firmulate.com/live.html provides a real-world look into the future of AI management: a landscape where trustworthiness, decision discipline, and resistance to manipulation are as vital as accuracy or speed. This is the next step toward AI that can not only do the work but also do it honestly — a must-have trait for safeguarding your organization’s integrity.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html