
Imagine a company that’s completely transparent about its struggles — an open window into AI managing real business risks, in real time.
At a time when health tech and wellness devices promise personalized insights, what if the same transparency extended to how AI actually performs in the messy, unpredictable world of business? Enter Firmulate, a groundbreaking experiment where artificial intelligence models run a small software company, facing the same crises, temptations, and decisions a human manager would encounter — but with every move recorded, analyzed, and made publicly accessible.
This isn’t just a demo or a chat interface. It’s a real, functioning enterprise that loses €105,000 each month against a modest €2,300 in recurring revenue. Yet, against all odds, it continues to operate, judged not only by its financials but by how well its AI managers navigate complex dilemmas under pressure.
The Live Company: Watching AI in Action
Accessible at firmulate.com/live.html, the live experiment shows the company in action — every workday, every decision, every crisis. There are 13 synthetic employees, each governed by a set of over 680 self-learned rules. These models are tasked with running the business as if they were human managers, and their decisions are fully auditable and versioned.
What makes this experiment particularly striking is its rigor: all models face the same week’s crises, customer demands, and temptations, and are tested against the same set of challenges. They’re given the same documents, the same scenarios, and are observed to see which can close deals, which refuse unethical shortcuts, and which falter under pressure.
Key Findings from the Experiment
- All four AI models successfully identified and responded to every crisis — a sign that they can recognize and react to business risks.
- Every model refused manipulation attempts, including social engineering tricks like fake CEO messages or reporter tricks — even when escalated over multiple stages.
- Despite this resilience, only two models managed to close a critical deal worth €55,000, bringing in extra monthly recurring revenue of +€4,583. One model, Opus 4.8, was the most thorough in analysis but left the close on the table, showing how discipline and focus matter just as much as intelligence.
- The decisive weakness was buried deep in the company’s own files — not in the customer interactions. The models that read and understood these internal documents won the deal at full price.
The Business Reality: A Money-Losing Machine
While the models demonstrate impressive decision-making capabilities, the company’s financials paint a stark picture — burning through €105,000 every month against a modest €2,300 in recurring revenue. A public cash countdown underscores the urgency of finding effective AI-driven management strategies — both for this experiment and for future real-world applications.
What the Experiments Reveal About AI in Business
The experiment underscores a crucial point: performance isn’t just about generating convincing chat or responses. It’s about execution — following through on decisions, reading internal documents thoroughly, and resisting unethical shortcuts under pressure.
For businesses considering AI automation, the takeaway is clear. The real challenge is not just in AI’s ability to analyze or generate language but in its capacity to finish what it starts, stay honest under temptations, and understand the full context of the organization.
Transparency and Build-in-Public Strategy
This experiment is a showcase of build-in-public in its most extreme form. Every decision, every rule learned, and every crisis response is publicly recorded and versioned. This level of transparency allows observers to learn not just about AI capabilities but also about the nuanced human-like failures that still occur — like missed opportunities or process slips, as seen with the Opus 4.8 model.
Interactive Engagement
Readers and enterprise clients can run their own test scenarios against a read-only export of their business at firmulate.com/pilot.html. This safe sandbox allows organizations to explore how their own decision-making can be challenged and improved by AI — without risking real systems or data.

Watching a real company managed entirely by AI models reveals that true business resilience depends not just on spotting crises but on decisively following through and resisting unethical shortcuts — all in full view. This experiment shows AI’s promise and limits, emphasizing that transparency and thoroughness are key to trustworthy automation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.