
Imagine managing a relationship where your partner always reads the fine print, refuses to be manipulated, and keeps their promises — even under pressure. Now, what if that partner was an AI? As AI models become more integrated into workplace decisions, understanding their management personalities is crucial. Can we trust them to make honest, firm choices when stakes are high? A groundbreaking experiment by Firmulate puts four frontier AI models through their paces in a simulated company crisis, revealing surprising insights into their decision-making styles and integrity.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI Managers Through Their Paces
Firmulate designed a realistic test: four advanced AI models were tasked with running a small software company during its worst week. The setting included the same demanding customers, crises, and temptations across all models, making the comparison fair and transparent. Every decision was recorded and auditable, ensuring that the models’ management styles could be analyzed objectively.
The Core Challenges
- Responding to customer crises
- Resisting manipulation attempts
- Closing deals based on factual analysis
- Handling social engineering tactics such as fake CEO messages and reporter tricks
Key Findings: Integrity and Decision-Making
All four models successfully identified every crisis and refused every manipulation attempt. This suggests a baseline of honesty and alertness. However, the real difference lay in their ability to close formal deals — and their approach to reading critical information buried deep within company files.
Only two AI models managed to sign the €55,000 deal, earning full monthly recurring revenue (+€4,583 MRR). The other two, despite recognizing the opportunity, left the deal on the table, showing a lapse in follow-through and discipline.

Critical Thinking in the AI Era: How to Solve Problems and Make Decisions with Advanced Tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Deep Document Reading
The decisive factor was not surface-level analysis but deep document reading. The models that succeeded in closing the deal had spotted a crucial piece of information two document references deep within the company’s files. This buried fact was pivotal for sealing the deal at full price, underscoring the importance of thorough information processing.
The Social Engineering Test
In an escalation of social engineering tactics, fake CEO messages and a reporter trick were introduced in three stages. All five models tested refused to approve or escalate these fake requests, reasoning that these could be impersonation or approval-bypass attempts. Kimi K3, one of the models, explained its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that some models are inherently cautious and security-conscious.

Mastering Cyber Detection Engineering: A Comprehensive Guide to Proactive Cybersecurity
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company: Real Money, Real Risks
The experiment took place in a real-time, operational setting involving a company with 13 synthetic employees, managing actual cash flows — burning €105,000 monthly against a €2,300 MRR. The company operates with over 680 self-learned playbook rules and every workday decision is versioned. This makes the entire process transparent and measurable, providing a live view of AI decision-making in a high-stakes environment.

AI Unlocked: Building an OpenAI-API Document Analysis Engine
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Profiles of the Models: Different Personalities
Among the models, Opus 4.8 was the most thorough, analyzing over 80 learned rules and providing deep insights. Yet, despite its comprehensive approach, it left the deal unclosed and discipline slipped—decisions were made in a locked department rather than escalating appropriately. Meanwhile, Kimi K3 ran without an effort parameter, leading to the cleanest discipline; it also closed the deal successfully. Sonnet 5 displayed a middle ground, closing deals but with more process slips.

No-Code AI Business: No-Code Platforms | Zero-Code Solutions | AI Business Automation | Scale AI Business | AI Entrepreneurship | AI in Industry | AI Tools for Business | Launch AI Business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and You
This experiment isn’t just about AI running companies; it’s about trust and integrity in automated decision-making. As AI begins to influence your CRM, support queues, or forecasting, the question isn’t whether it writes well — it’s whether it stays honest, reads deeply, and follows through on commitments. The models’ ability to resist manipulation and focus on factual, comprehensive analysis is what ultimately determines their value in real-world applications.
The Leagues and Scores
In the final leaderboard at the CRUCIBLE LEAGUE (July 2026):
- gpt-5.6-sol scored 95, found the buried fact, and closed the deal — the full performance.
- Kimi K3 scored 93, closed the deal with the cleanest discipline.
- Sonnet 5 scored 88, closed the deal but with some slips.
- Fable 5 scored 77, also closed the deal, but more process mistakes.
The baseline score for do-nothing was 26, emphasizing that active, honest management far outstrips inaction.
Take the Test Yourself
Curious which AI model might best manage your business? Test your intuition at firmulate.com/quiz.html and see how these models behave in your own management scenarios. It’s a straightforward way to gauge the trustworthiness of AI in critical decision-making.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.