
Imagine an AI that doesn’t just chat but reads your company’s internal documents before making a decision—potentially saving or losing millions. In a groundbreaking live experiment, AI models faced off in a simulated business crisis, revealing a surprising truth: the difference between winning and losing a €55,000 deal was buried two documents deep in company files. This effort to see what AI truly understands could reshape how businesses trust and deploy artificial intelligence.
The Experiment: Testing AI in Real-World Business Scenarios
Firmulate’s latest live experiment placed four advanced AI models in a realistic simulation of a small software company’s worst week. The models had to navigate the same customer crises, internal dilemmas, and manipulative tactics—think fake CEO messages and reporter tricks—just like a real management team would. Every decision was captured and made auditable, providing a clear window into their reasoning and behavior.
enterprise AI document reading software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Key Find: Reading Deep Matters
While all four AI models successfully identified and refused every crisis and manipulation attempt, only two managed to close the deal worth €55,000—an outcome linked to a crucial detail buried two references deep in the company’s files, not in the customer-facing information. The models that read and understood this buried fact won the contract at full price, worth an additional €4,583 in monthly recurring revenue (MRR).
What This Means for Business AI
This experiment underscores an essential insight: for AI to truly perform as a trustworthy business partner, it must look deeper than surface level. The models that read and interpret internal documents before responding made better decisions, especially in high-stakes scenarios. Conversely, models that relied only on superficial data or that didn’t dig deep risk missing critical context, leading to missed opportunities or mistakes.
Wider Implications: Trust and Integrity Under Pressure
The models were tested against social engineering attempts, including staged messages from a fake CEO and background checks by reporters. All refused to be manipulated, but only those with thorough understanding and disciplined processes secured the deal. The takeaway?: Trustworthiness isn’t just about avoiding scams; it’s about having the discipline and insight to recognize what’s important—especially when it’s buried in the weeds.
The Live Business: A Real-World Simulation with Live Money
The experiment isn’t just theoretical. The company in question, with 13 synthetic employees managing real money mechanics, burns €105,000 monthly against a revenue of €2,300. Every day, the AI models are tested in a live environment, with over 680 self-learned rules guiding their behavior, making this a transparent, watchable testbed for enterprise AI performance.
What Different AI Models Showed
The models scored differently. The most thorough, Opus 4.8, with over 80 learned rules and deep analysis, placed last—not because it failed but because discipline slipped, and it left important opportunities on the table. Meanwhile, Kimi K3, running without an effort parameter, closed the deal with clean discipline, showcasing that focus and default settings influence outcomes.
Why You Should Care
If AI agents will manage your customer data, support queues, or forecasting, the crucial question isn’t just about how well they generate chatter. It’s whether they finish what they start, read your internal files thoroughly, stay honest under pressure, and understand the full context before acting. This experiment proves that superficial performance isn’t enough—deep understanding is key to trustworthy, effective AI.
Next Steps: Wargaming Your Own Business
For enterprises eager to test AI readiness, Firmulate offers a pilot program allowing you to run your own simulated crises with your data—without risking real systems. This ‘wargame’ reveals where your AI might slip and helps you prepare for deployment that’s both effective and trustworthy.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html