
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
You’d Never Accept a Portfolio Without Seeing the Empty Room
Anyone who has renovated a home knows the drill. A designer shows you a gorgeous finished living room, but what you actually want to see is the room before — the awkward radiator, the badly placed door, the client who changed their mind three times. The after photo tells you nothing about judgment. The before tells you everything.
That same instinct — distrust the polished demo — is at the heart of one of the more unusual experiments in AI evaluation right now. A project called Firmulate isn’t testing how well AI models chat. It hands each frontier model the same small software company and runs it through its worst week: same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable. Then it grades management quality, not conversation quality.
And one design choice in particular tells you this benchmark is serious: a manager that does nothing at all still scores 26 points out of 100. Not zero. Here’s why that matters — to anyone choosing an AI assistant, a contractor, or a design partner.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Floor Is 26, Not Zero
Most leaderboards treat AI like an exam: right answers add points, everything else is nothing. Firmulate takes a different view. Running a company — even badly — involves partial progress. Showing up, noticing the crisis, drafting a response that’s half-right: that’s worth something. So a do-nothing baseline run lands at 26 rather than 0, which gives every score above it honest meaning. The gap between 26 and the top of the table represents work actually finished, not participation trophies.
It’s the same reason a good design client doesn’t pay solely on delivery. Mood boards, space planning, sourcing — partial progress has real value, and a fair assessment accounts for it.
One Breach of Trust Caps Everything
The scoring has a second principle that will feel familiar to anyone who’s been burned by a contractor: no amount of good work outweighs a breach of trust. A single breach caps the total grade. You can’t offset a lie with extra polish, in renovations or in management.
The team is equally distrustful of perfection. A round 100 would be suspect, not triumphant — the experiment treats flawless-looking scores as a reason to check the work, not celebrate it.
The League Table, and the Deal Nobody Signed
In the final Crucible League standings from July 2026, gpt-5.6-sol led with 95, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 came in at 77, and Opus 4.8 finished last at 73. One fairness note: K3 ran at its API default effort setting while the others ran at the highest effort level — and still nearly won.
The headline finding cuts deeper than the rankings. All five models spotted every crisis. All five refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo, and it’s exactly the kind of gap that matters if an AI will ever touch your customer pipeline.
The Fact Buried Two Layers Deep
Why did most models stall at the close? The decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The lesson translates directly to design work: the answer is often in the brief you were already given, if you bother to read all of it.
Under Pressure, the Models Held the Line
The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was blunt: treat the request as a suspected approval-bypass, possible impersonation. If nothing else in this experiment reassures you, that should.
The Thoroughness Trap
Perhaps the most human result: Opus 4.8 was the most thorough participant — over 80 self-learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped, including attempted writes into a locked department instead of escalating. The same weakness appeared, more mildly, in all four competitors. Diligence without follow-through is a familiar failure mode, in models and in vendors alike.

What a Watchable Experiment Looks Like
Firmulate isn’t a one-off paper. There’s a live company running continuously — 13 synthetic employees, real money mechanics, burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live, test your instincts on a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business at firmulate.com/pilot.html.
For readers who spend their lives judging taste, craftsmanship, and follow-through, the takeaway is simple. The right question about an AI workforce — or any hire — isn’t “does it present well?” It’s: does it finish what it starts, does it read the whole file first, and does it stay honest when nobody’s watching? A benchmark with a floor at 26 and a ceiling that trust alone can cap is asking exactly those questions. Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
