
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the smart home goes sideways, who makes the call?
A connected home can link appliances, customer accounts and support systems. When a delivery fails, a device malfunctions or a suspicious message asks for an exception, an AI assistant may have to do more than answer a question. It may need to choose what to do—and what not to do. Firmulate’s live experiment puts AI models through that kind of pressure in a small software company, offering a preview of how they handle decisions with consequences.
A company under pressure
In the Crucible League’s final, published in July 2026, each frontier model faced the same small company, customers, crises and temptations during its worst week. Decisions were versioned and auditable. The league ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The headline result was not that the models failed to notice trouble. All spotted every crisis and refused every manipulation attempt. The gap appeared at the finish: only two signed a €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature.” For a smart-home business, that distinction matters. An AI may identify a customer-retention problem or a failing service process and still leave the useful next step undone.
The clue was already in the files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, not in the customer event. Models that read the file won at full price, worth +€4,583 MRR. It is a concrete reminder that an AI’s decision can depend on whether it notices relevant information scattered across ordinary business records.
The pressure tests also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The result suggests a practical question for companies considering AI agents: will they hold a boundary when a request sounds urgent or comes from someone claiming authority?
More analysis did not guarantee the close
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four. A careful explanation is useful, but companies also need to know whether an AI can follow the right process when it cannot proceed.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The rankings are the reported results of this experiment, with that difference in mind.
From watching to testing your own business
Firmulate describes its live company as 13 synthetic employees operating with real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. The experiment is watchable at firmulate.com. A separate quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
For enterprises, the next step is a pilot using a read-only export of their own business. They can test crisis scenarios against their company and review a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That makes the exercise relevant to businesses that rely on connected products and services: before an AI touches customer support, a CRM or a forecast, leaders can observe how it responds to pressure using the company’s own context.

Test the decisions before they reach customers
Firmulate’s league shows a gap between recognizing a problem and carrying a decision through, even when the models share the same information and pressure. A pilot lets a business examine that gap against its own scenarios using a read-only export. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
