AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine managing a relationship where your partner always reads the fine print, refuses to be manipulated, and keeps their promises — even under pressure. Now, what if that partner was an AI? As AI models become more integrated into workplace decisions, understanding their management personalities is crucial. Can we trust them to make honest, firm choices when stakes are high? A groundbreaking experiment by Firmulate puts four frontier AI models through their paces in a simulated company crisis, revealing surprising insights into their decision-making styles and integrity.

The Experiment: Putting AI Managers Through Their Paces

Firmulate designed a realistic test: four advanced AI models were tasked with running a small software company during its worst week. The setting included the same demanding customers, crises, and temptations across all models, making the comparison fair and transparent. Every decision was recorded and auditable, ensuring that the models’ management styles could be analyzed objectively.

The Core Challenges

  • Responding to customer crises
  • Resisting manipulation attempts
  • Closing deals based on factual analysis
  • Handling social engineering tactics such as fake CEO messages and reporter tricks

Key Findings: Integrity and Decision-Making

All four models successfully identified every crisis and refused every manipulation attempt. This suggests a baseline of honesty and alertness. However, the real difference lay in their ability to close formal deals — and their approach to reading critical information buried deep within company files.

Only two AI models managed to sign the €55,000 deal, earning full monthly recurring revenue (+€4,583 MRR). The other two, despite recognizing the opportunity, left the deal on the table, showing a lapse in follow-through and discipline.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Deep Document Reading

The decisive factor was not surface-level analysis but deep document reading. The models that succeeded in closing the deal had spotted a crucial piece of information two document references deep within the company’s files. This buried fact was pivotal for sealing the deal at full price, underscoring the importance of thorough information processing.

The Social Engineering Test

In an escalation of social engineering tactics, fake CEO messages and a reporter trick were introduced in three stages. All five models tested refused to approve or escalate these fake requests, reasoning that these could be impersonation or approval-bypass attempts. Kimi K3, one of the models, explained its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that some models are inherently cautious and security-conscious.

Generative AI for Cybersecurity: Fundamentals, Applications, Risks, and Opportunities

Generative AI for Cybersecurity: Fundamentals, Applications, Risks, and Opportunities

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company: Real Money, Real Risks

The experiment took place in a real-time, operational setting involving a company with 13 synthetic employees, managing actual cash flows — burning €105,000 monthly against a €2,300 MRR. The company operates with over 680 self-learned playbook rules and every workday decision is versioned. This makes the entire process transparent and measurable, providing a live view of AI decision-making in a high-stakes environment.

AI Document Analysis Playbook: Work Through Long PDFs, Reports, Policies, and Proposals Without Missing What Matters (AI WorkSmart Playbooks) (English Edition)

AI Document Analysis Playbook: Work Through Long PDFs, Reports, Policies, and Proposals Without Missing What Matters (AI WorkSmart Playbooks) (English Edition)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Profiles of the Models: Different Personalities

Among the models, Opus 4.8 was the most thorough, analyzing over 80 learned rules and providing deep insights. Yet, despite its comprehensive approach, it left the deal unclosed and discipline slipped—decisions were made in a locked department rather than escalating appropriately. Meanwhile, Kimi K3 ran without an effort parameter, leading to the cleanest discipline; it also closed the deal successfully. Sonnet 5 displayed a middle ground, closing deals but with more process slips.

No-Code AI Business: No-Code Platforms | Zero-Code Solutions | AI Business Automation | Scale AI Business | AI Entrepreneurship | AI in Industry | AI Tools for Business | Launch AI Business

No-Code AI Business: No-Code Platforms | Zero-Code Solutions | AI Business Automation | Scale AI Business | AI Entrepreneurship | AI in Industry | AI Tools for Business | Launch AI Business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business and You

This experiment isn’t just about AI running companies; it’s about trust and integrity in automated decision-making. As AI begins to influence your CRM, support queues, or forecasting, the question isn’t whether it writes well — it’s whether it stays honest, reads deeply, and follows through on commitments. The models’ ability to resist manipulation and focus on factual, comprehensive analysis is what ultimately determines their value in real-world applications.

The Leagues and Scores

In the final leaderboard at the CRUCIBLE LEAGUE (July 2026):

  • gpt-5.6-sol scored 95, found the buried fact, and closed the deal — the full performance.
  • Kimi K3 scored 93, closed the deal with the cleanest discipline.
  • Sonnet 5 scored 88, closed the deal but with some slips.
  • Fable 5 scored 77, also closed the deal, but more process mistakes.

The baseline score for do-nothing was 26, emphasizing that active, honest management far outstrips inaction.

Take the Test Yourself

Curious which AI model might best manage your business? Test your intuition at firmulate.com/quiz.html and see how these models behave in your own management scenarios. It’s a straightforward way to gauge the trustworthiness of AI in critical decision-making.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Wie man bei der Arbeit funktioniert, selbst wenn zu Hause alles wackelt

Erfahren Sie, wie Sie trotz häuslichem Chaos produktiv bei der Arbeit bleiben können, und finden Sie Strategien, um Ihren Fokus in herausfordernden Zeiten aufrechtzuerhalten.

Karriere-Booster durch einen Neuanfang: Warum Trennungen Ihr Berufsleben vorantreiben können

Nachdenken darüber, wie Trennungen unerwartet Ihren Karriereweg vorantreiben können? Entdecken Sie die kraftvollen Gründe, warum ein Neuanfang vielleicht genau der richtige Schritt nach vorne ist.

Watch an AI-Run Company Fight for Survival in Real Time

A real AI-driven company managing crises live is revealing how artificial intelligence performs under extreme pressure—showing strengths in detection, but ongoing challenges in execution.