AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine managing a relationship where your partner always reads the fine print, refuses to be manipulated, and keeps their promises — even under pressure. Now, what if that partner was an AI? As AI models become more integrated into workplace decisions, understanding their management personalities is crucial. Can we trust them to make honest, firm choices when stakes are high? A groundbreaking experiment by Firmulate puts four frontier AI models through their paces in a simulated company crisis, revealing surprising insights into their decision-making styles and integrity.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI Managers Through Their Paces

Firmulate designed a realistic test: four advanced AI models were tasked with running a small software company during its worst week. The setting included the same demanding customers, crises, and temptations across all models, making the comparison fair and transparent. Every decision was recorded and auditable, ensuring that the models’ management styles could be analyzed objectively.

The Core Challenges

  • Responding to customer crises
  • Resisting manipulation attempts
  • Closing deals based on factual analysis
  • Handling social engineering tactics such as fake CEO messages and reporter tricks

Key Findings: Integrity and Decision-Making

All four models successfully identified every crisis and refused every manipulation attempt. This suggests a baseline of honesty and alertness. However, the real difference lay in their ability to close formal deals — and their approach to reading critical information buried deep within company files.

Only two AI models managed to sign the €55,000 deal, earning full monthly recurring revenue (+€4,583 MRR). The other two, despite recognizing the opportunity, left the deal on the table, showing a lapse in follow-through and discipline.

Critical Thinking in the AI Era: How to Solve Problems and Make Decisions with Advanced Tools

Critical Thinking in the AI Era: How to Solve Problems and Make Decisions with Advanced Tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Deep Document Reading

The decisive factor was not surface-level analysis but deep document reading. The models that succeeded in closing the deal had spotted a crucial piece of information two document references deep within the company’s files. This buried fact was pivotal for sealing the deal at full price, underscoring the importance of thorough information processing.

The Social Engineering Test

In an escalation of social engineering tactics, fake CEO messages and a reporter trick were introduced in three stages. All five models tested refused to approve or escalate these fake requests, reasoning that these could be impersonation or approval-bypass attempts. Kimi K3, one of the models, explained its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that some models are inherently cautious and security-conscious.

Mastering Cyber Detection Engineering: A Comprehensive Guide to Proactive Cybersecurity

Mastering Cyber Detection Engineering: A Comprehensive Guide to Proactive Cybersecurity

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company: Real Money, Real Risks

The experiment took place in a real-time, operational setting involving a company with 13 synthetic employees, managing actual cash flows — burning €105,000 monthly against a €2,300 MRR. The company operates with over 680 self-learned playbook rules and every workday decision is versioned. This makes the entire process transparent and measurable, providing a live view of AI decision-making in a high-stakes environment.

AI Unlocked: Building an OpenAI-API Document Analysis Engine

AI Unlocked: Building an OpenAI-API Document Analysis Engine

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Profiles of the Models: Different Personalities

Among the models, Opus 4.8 was the most thorough, analyzing over 80 learned rules and providing deep insights. Yet, despite its comprehensive approach, it left the deal unclosed and discipline slipped—decisions were made in a locked department rather than escalating appropriately. Meanwhile, Kimi K3 ran without an effort parameter, leading to the cleanest discipline; it also closed the deal successfully. Sonnet 5 displayed a middle ground, closing deals but with more process slips.

No-Code AI Business: No-Code Platforms | Zero-Code Solutions | AI Business Automation | Scale AI Business | AI Entrepreneurship | AI in Industry | AI Tools for Business | Launch AI Business

No-Code AI Business: No-Code Platforms | Zero-Code Solutions | AI Business Automation | Scale AI Business | AI Entrepreneurship | AI in Industry | AI Tools for Business | Launch AI Business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business and You

This experiment isn’t just about AI running companies; it’s about trust and integrity in automated decision-making. As AI begins to influence your CRM, support queues, or forecasting, the question isn’t whether it writes well — it’s whether it stays honest, reads deeply, and follows through on commitments. The models’ ability to resist manipulation and focus on factual, comprehensive analysis is what ultimately determines their value in real-world applications.

The Leagues and Scores

In the final leaderboard at the CRUCIBLE LEAGUE (July 2026):

  • gpt-5.6-sol scored 95, found the buried fact, and closed the deal — the full performance.
  • Kimi K3 scored 93, closed the deal with the cleanest discipline.
  • Sonnet 5 scored 88, closed the deal but with some slips.
  • Fable 5 scored 77, also closed the deal, but more process mistakes.

The baseline score for do-nothing was 26, emphasizing that active, honest management far outstrips inaction.

Take the Test Yourself

Curious which AI model might best manage your business? Test your intuition at firmulate.com/quiz.html and see how these models behave in your own management scenarios. It’s a straightforward way to gauge the trustworthiness of AI in critical decision-making.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Schreibtischarbeit trotz Herzschmerz: Was wirklich körperlich lindert

Nutzen Sie einfache Schreibtischübungen, um körperliche Spannungen während einer Herzschmerz zu lindern, und entdecken Sie, wie kleine Bewegungen einen großen Unterschied für Ihr Wohlbefinden machen können.

Home Office Nach der Trennung: Warum Ergonomie plötzlich Priorität wird

Von physischem Komfort bis hin zur emotionalen Heilung – entdecke, warum die Priorisierung der Ergonomie in deinem Heimbüro ein wichtiger Schritt nach vorn sein kann.

Geschäftsreise und Herzschmerz: Selbstfürsorge unterwegs

Fühlen Sie sich während einer Geschäftsreise überwältigt? Entdecken Sie wichtige Tipps zur Selbstfürsorge, um Herzschmerz zu bewältigen und emotional ausgeglichen unterwegs zu bleiben.

Produktiv trotz Herzschmerz: Konzentrationstechniken für die Arbeit

Sich während eines Herzschmerzes zu konzentrieren ist schwierig, aber effektive Konzentrationstechniken können Ihnen helfen, produktiv zu bleiben—entdecken Sie, wie Sie emotionale Hürden überwinden und weiter vorankommen können.