
Imagine trusting a new partner in a high-stakes game — one who promises honesty, resilience, and sharp decision-making. Would you bet on the newcomer over seasoned players? Recent experiments suggest that in the world of AI-driven business management, fresh entrants are proving they can outperform the established giants, even in their toughest moments.
The Fight for AI Leadership in Business Management
In a groundbreaking live experiment, four advanced AI models were pitted against each other by managing a real, functioning small software company. This wasn’t a typical test — it was a simulation of the company’s worst week, packed with crises, customer demands, and ethical temptations. The goal was simple: see which AI could navigate the storm and close a major deal worth €55,000, all while maintaining integrity and discipline.

Building AI-Powered Products: The Essential Guide to AI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League of Leading AI Models
The results speak volumes. At the top was gpt-5.6-sol, scoring 95 points, just ahead of the newcomer Kimi K3 with 93 points. Not far behind was Sonnet 5 at 88, then Fable 5 at 77, and Opus 4.8 at 73. These scores are from a final leaderboard in July 2026, reflecting how well each model performed in this intense, real-world scenario.

AI for Real Companies: A Practical Guide to Smarter Systems and Stronger Profits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decisive Factors: Reading Files and Integrity
While all models detected every crisis and refused every manipulation attempt — including a staged ‘fake CEO’ message — the key difference lay in their ability to uncover buried information in the company’s own files. The most successful model, Kimi K3, found a crucial reference deep in the company’s documentation, enabling it to close the deal at full price, adding €4,583 in monthly recurring revenue.
In contrast, Opus 4.8, despite being the most thorough participant with over 80 learned rules and deep analysis, left some opportunities unseized and slipped in discipline, such as failing to escalate issues properly. Interestingly, Opus ran with a different setup, with no effort parameter, which may have influenced its performance.

Ethics and Integrity in Education (Practice): Derived from the 9th European Conference on Ethics and Integrity in Academia (Ethics and Integrity in Educational Contexts, Band 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Ethics Under Pressure
All four models demonstrated resilience against social engineering attempts, refusing to sign off on fake CEO requests or accept dubious background cues — a crucial aspect for real-world deployment. Kimi K3’s on-record reasoning emphasized treating suspicious requests as potential impersonation attempts, a stance that likely contributed to its success.

Executive Cyber Crisis & Board Tabletop Exercise Workbook: 12 Ready-to-Run CEO, Board, Ransomware, Data Breach, AI, Vendor & Reputation Crisis Simulations with Timed Injects, Decision Cards
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Company Behind the Experiment
The live simulation involved a company with 13 synthetic employees, managing real money mechanics with a burn rate of €105,000 monthly against a modest €2,300 MRR. The system is transparent, versioned daily, and accessible at firmulate.com/live. It’s designed to test whether AI can handle critical business decisions, stay honest, and deliver real value — not just chatty responses.
Implications for Business and Beyond
This experiment underscores that the question is no longer whether AI can generate convincing dialogue but whether it can reliably complete complex, integrity-critical tasks. With the league open and newcomers like Kimi K3 proving their mettle, business leaders face a new reality: selecting an AI isn’t just about scores or buzzwords but about actual performance under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html