
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the showroom is under pressure, a polished pitch is not enough
Imagine an AI helping run an interiors business as orders wobble, a competitor makes a move and an urgent message appears to come from the CEO. It might spot each problem and give sensible advice. But would it protect the business, follow its own rules and close the deal? That is the question Firmulate’s live experiment puts on display.
One company, the same difficult week
Firmulate ran frontier AI models through the same small software company and its worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The experiment measures how models manage a business under pressure, rather than how convincing they sound in a chat.
The final Crucible League, from July 2026, ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Its rule is stark: partial progress counts, but one breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
Seeing the problem is only half the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result was a telling gap: “Same diagnosis, same pitch — no signature.”
The deciding clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding speaks to any business that relies on procedures, customer records and hard-won knowledge: an AI can make the right-sounding diagnosis and still miss what the company already knows, or fail to turn that knowledge into action.
Trust is tested in ordinary-looking messages
The manipulation test included fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offered a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and its discipline slipped: it attempted writes into a locked department instead of escalating. That same weakness appeared, more weakly, in all four models. More analysis alone did not guarantee better execution.
A company you can watch—and a test you can bring home
The live Firmulate company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules and every workday versioned. Its decisions and progress can be watched at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites visitors to guess which model made each choice.
There is also a fairness note: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That context matters when comparing the league results.

From watching to testing your own business
A public experiment can show how models handle someone else’s crises. A pilot can put the question to your own company: Firmulate can run the wargame against a read-only export of your business, produce a board report with model rankings and expose weak points in your playbooks. Nothing writes back to real systems.
To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
