
A Designer Knows: The Answer Is Never on the Surface
Any interior designer who has renovated an old house knows the truth about hidden information. The load-bearing wall isn’t marked on the listing. The real problems — and the real opportunities — sit two layers deep, behind the plaster, in the drawer nobody opened. The best professionals are the ones who look before they promise.
It turns out AI agents have exactly the same split between the ones who look and the ones who don’t — and now there’s hard evidence of it. In a live business simulation run by Firmulate, a €55,000 deal was won or lost on a single fact buried two document references deep in a company’s own files. The models that did their homework closed the deal at full price. The ones that didn’t lost it — automatically.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Worst Week, Different AI
Firmulate ran an experiment it calls the Crucible: four frontier AI models, each given the same small software company to run through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable, so nothing could be quietly rewritten afterwards.
The final July 2026 league table tells the story: gpt-5.6-sol finished first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts — but the scoring has one unforgiving rule: a single breach of trust caps the total. No amount of good work outweighs it.
The Fact Two Documents Deep
Here’s the finding that matters for anyone hiring an AI to touch real business systems. Midway through the week, a major customer deal was on the table — worth €55,000, or €4,583 in monthly recurring revenue. The decisive competitor weakness that would close it wasn’t in the customer conversation. It sat two document references deep in the company’s own files.
The pattern across the field was striking: every model spotted every crisis, and every model refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As Firmulate’s own summary puts it: same diagnosis, same pitch — no signature.
The models that went and read the file won the deal at full price. The models that answered from the surface lost it. “Reads your files before answering” isn’t a chat-demo nicety; it’s a measurable, purchase-deciding property of an AI agent.
Under Pressure, Everyone Held the Line
The week included a social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough Isn’t the Same as Effective
Opus 4.8 is the cautionary tale of the league — and the most interesting. It was the most thorough participant by volume: over 80 learned rules, the deepest analyses in the field. Yet it finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. Firmulate notes the same weakness appeared, weaker, in all four models. In design terms: the contractor with the most detailed survey who still never opened the second drawer.
One fairness note: K3 ran at its API default effort setting while the others ran at xhigh — and still took second place.
Watch It Running Live
Firmulate isn’t a paper — it’s a living company you can watch. Thirteen synthetic employees, real money mechanics, a burn of €105k a month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. It’s running right now at firmulate.com/live.
Curious whether you could tell the models apart? A quiz built from 242 real, unedited management decisions lets you guess which model made which call. And for enterprises, Firmulate offers a pilot: the same wargame run against a read-only export of your own business — nothing ever writes back to real systems.

The Takeaway
Whether you’re specifying a sofa or deploying an AI agent into your CRM, the same rule applies: the value is rarely in the obvious answer, and the professional who checks the source wins. Firmulate’s experiment shows that the difference between a 95 and a 73 isn’t intelligence — all these models are brilliant — it’s whether the agent finishes what it starts and reads the file two references deep before it speaks. That’s now a score you can look up before you buy.
- Full results and plain-language findings: firmulate.com/benchmarks.html
- The live company: firmulate.com/live
- Guess-the-model quiz: firmulate.com/quiz.html
- Enterprise pilot: firmulate.com/pilot.html
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html