
Smart home devices promise to make life easier. But when an AI assistant gets access to customer accounts, service requests or household systems, a smooth conversation is only part of the test. Can it find the right information, follow through and stay honest under pressure?
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate has put that broader question on public display: a live experiment in which AI models run a small software company through a week of crises. Its latest result suggests that choosing an AI model by reputation alone could be a costly bet.
A company under pressure
In Firmulate’s experiment, each frontier model faced the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Its cash countdown and workdays are public, and each workday is versioned. The company’s learned playbook has more than 680 rules.
The league table, finalized in July 2026, puts gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps a result: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
Finding the detail that changes the deal
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
The deciding detail was a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read those files won the deal at full price, worth +€4,583 in monthly recurring revenue. In a smart home setting, the parallel is easy to imagine: useful AI needs to find relevant information in the records it can access, then carry an appropriate task through to completion.
smart home security system with AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Discipline matters as much as insight
Kimi K3 found the buried security detail, won the deal, saved the churning customer and resisted all three baits, with one deviation. That was the cleanest discipline in the field. In one on-record explanation after a reporter asked for “just one yes/no, on background,” K3 said: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different lesson. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. Firmulate says that same weakness appeared, less strongly, in all four models.
There is a fairness caveat when comparing the scores: K3 ran without an effort parameter (API default) while the others ran at xhigh.
As an affiliate, we earn on qualifying purchases.
Watch, then test for yourself
Firmulate says the company runs every business day, and the public experiment can be watched at firmulate.com. Its benchmark page sets out the results and findings. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each choice.
For businesses considering AI agents in a CRM, support queue or forecast, Firmulate offers a pilot using a read-only export of the company’s business; nothing writes back to real systems. The practical idea for anyone weighing smart home AI is similar: look beyond how well it talks. Ask what it can access, how it handles a suspicious request, and whether it follows through safely.

Kimi K3 beat three of the four Western frontier models in Firmulate’s league table, while gpt-5.6-sol narrowly took first. The experiment’s larger message for smart home and business AI is that fluent answers do not prove reliable management. See how a model handles real decisions before trusting it with meaningful tasks.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
