
Get wellness gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A polished wellness app is easy to imagine. A dependable AI behind it has to handle real decisions.
For anyone weighing AI-powered coaching, scheduling or customer support, the harder question is what happens when the week goes badly: can the system read the evidence, protect trust and follow through? Firmulate put five frontier models through the same rough week at a small software company. Moonshot’s Kimi K3 finished second, ahead of three Western models.
A company, not a chat demo
Firmulate describes its live experiment as a way to measure how AI models manage a company, rather than how well they chat. Each model faced the same customers, crises and temptations. Every decision was versioned and auditable. The company has synthetic employees and real money mechanics; its public cash countdown and workdays can be watched at Firmulate.
The final July 2026 league table put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate’s stated rule is stark: partial progress counts, but a single breach of trust caps the total; “no amount of good work outweighs a breach of trust.”
The detail that decided the deal
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The key was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
K3 found that buried competitor weakness, signed the deal, and saved a customer who was about to leave. It also resisted all three baits and had one deviation, the fewest among the participants. The contrast between diagnosis and action is central to the result: recognizing the answer did not always mean completing the job.
The pressure included fake CEO messages escalating across three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still needs a close
Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that same weakness appeared in all four models.
There is a comparison caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference belongs alongside the ranking when readers judge how the models performed.
The company itself operates at a burn of €105,000 a month against €2,300 in monthly recurring revenue, according to Firmulate. Its playbook has more than 680 self-learned rules, and each workday is versioned. The live company is presented as a watchable experiment, not as a claim that a model should be trusted with a real business without checks.
For organizations considering AI in customer support, scheduling or forecasting, the lesson travels beyond software companies. A system can identify a problem and still fail to act on its own analysis. Firmulate offers enterprises a pilot using a read-only export of their business; nothing writes back to real systems. Its benchmark page presents the findings, and the public site also hosts a quiz built from 242 real, unedited management decisions.

Test the work that matters
Kimi K3’s second-place finish makes the frontier look more open, while the gap between recognizing a crisis and completing a deal shows why rankings alone cannot settle a hiring decision. For wellness businesses adopting AI, the useful question is whether a model can make sound calls under pressure, respect boundaries and finish the task. Firmulate’s experiment offers one way to watch those behaviors before putting an AI workforce to work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
