
Imagine having a virtual boss that not only makes decisions but also exhibits a distinct personality—whether it’s terse, thorough, or cautious. As businesses increasingly rely on AI for management, understanding these virtual personalities might be the key to smarter automation. But can we really tell one AI apart from another based on their decisions? The live experiment by Firmulate offers a revealing look.
The Experiment: Putting AI Models to the Test in a Live Business Crisis
In a groundbreaking real-world test, four frontier AI models were challenged to run a small software company through its most tumultuous week. They faced the same customers, the same crises, and the same temptations to cut corners. Every decision was carefully tracked and made transparent, providing an unprecedented view into how different AI models behave when managing real business operations.
Who Were the Competitors?
- GPT-5.6-sol 95: The top scorer, which not only identified the company’s hidden information but also successfully closed a €55,000 deal, marking full achievement.
- Kimi K3 93: A newcomer that squeezed out a win with the cleanest decision-making and highest discipline.
- Sonnet 5 88: Managed to close the deal but with some slips in process discipline.
- Fable 5 77: Also closed the deal but showed more signs of wavering under pressure.
Interestingly, the baseline—an AI that does nothing—scored just 26, emphasizing how much better these models perform with proper decision-making.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Insights: More Than Just Chat Quality
All four models successfully detected and refused manipulative social engineering attempts, like staged CEO messages and reporter tricks. For instance, K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that these models aren’t just surface-level chatbots—they understand context and risks.
The critical difference lay in their depth of analysis. The leading model, GPT-5.6-sol 95, read deeply into internal company documents and uncovered a crucial piece of information that sealed the deal at full price. Models that failed to read those documents lost potential revenue, illustrating that reading comprehension and internal knowledge access matter immensely.
The Hidden Weaknesses and Discipline Gaps
While Opus 4.8 ran the most rules (over 80) and performed thorough analyses, it ultimately left the close opportunity on the table, showing a discipline slip that cost it the win. Despite the detailed analysis, it failed to escalate critical decisions to the right channels, revealing that even the most comprehensive AI can falter in operational discipline under stress.
What Does This Mean for Businesses?
This experiment isn’t just about AI wizardry; it’s about how these models behave when managing real-world, profit-critical tasks. The findings are clear:
- AI models can detect crises and refuse unethical shortcuts, demonstrating integrity.
- Deeper understanding and document reading capabilities lead to better outcomes.
- Discipline and decision discipline are crucial—more rules don’t guarantee success, especially if critical decisions are not escalated properly.
For companies contemplating AI management assistants, these results suggest that choosing the right model isn’t just about raw scores but understanding their management style and decision discipline. The models with higher scores also performed better at reading internal information and closing deals at full price—valuable traits for any enterprise.
Experience It Live
Curious to see these models in action? The same company’s live operations run every business day and are openly accessible. You can watch real decisions, read actual employee communications, and even run your own game against a read-only version of your business data. Visit firmulate.com/live and see how AI manages in real time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html