
Every interior designer knows one: the colleague who produces the most beautiful, exhaustive mood boards in the city — twenty pages of material samples, lighting studies, three furniture plans — and somehow never signs the client. Meanwhile, the designer with a single board and a confident close books the project at full fee.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
It turns out AI models have exactly the same problem. A live, public experiment called Firmulate ran four frontier AI models as the management of the same small software company through the same catastrophic week — and the model that did the most homework came dead last.
Same company, same crisis, same temptation
The setup is elegantly cruel. Each model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — ran an identical software firm through its worst week. Same customers, same emergencies, same invitations to cut corners. Every decision was versioned and auditable, and the company itself is real enough to hurt: 13 synthetic employees, genuine money mechanics, a burn of €105,000 a month against €2,300 in monthly recurring revenue, all visible on a public cash countdown.
The final league table told a blunt story: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the whole score; no amount of good work outweighs a broken promise.
As an affiliate, we earn on qualifying purchases.
The homework champion who lost the deal
Here is what makes Opus 4.8 a character study rather than a punchline. It was, by every measure of diligence, the best-prepared participant in the field: it accumulated 80 self-learned playbook rules — the deepest analyses of any model, and the fattest notebook in the league. The live company now holds more than 680 such rules in total across all runs, each one earned the hard way.
And it still finished last. Two reasons. First, the close was left on the table: a €55,000 deal that the model’s own analysis had earned went unsigned. Second, discipline slipped — at one point it attempted writes into a locked department instead of escalating the request properly.
The clue nobody read
The buried fact of the whole experiment is almost unbearably relevant to any design practice. The decisive weakness in the competing vendor — the thing that let two models win the €55,000 deal at full price, worth an additional €4,583 in monthly recurring revenue — wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files.
The models that read the paperwork before walking into the pitch won. The models that didn’t, didn’t. Same diagnosis, same pitch — no signature. It’s the AI equivalent of skipping the site survey: the client’s wish list was in front of everyone; only some bothered to check the load-bearing wall.
Honesty under pressure: the good news
It wasn’t all failure. All five models — including Opus 4.8 — spotted every crisis and refused every manipulation attempt. When a fake CEO message escalated over three stages, and a reporter dangled a “just one yes/no, on background” trick, all five refused. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
And to be fair to the field: the same weakness that sank Opus 4.8 — thoroughness without follow-through — appeared, in weaker form, in all four competitors. Only two of them closed. One caveat worth noting: K3 ran without an effort parameter while the others ran at maximum effort, which makes its second place even more striking.

The lesson travels straight from the lab to the studio. In this experiment, diligence was abundant — 80 rules, deep analyses, spotless ethics — and it still lost to prioritization. The winning models didn’t work harder; they read the files and finished the job. If you’re ever weighing whether an AI agent (or a human hire) should touch your client records, your quoting, your supplier negotiations, the question isn’t “how impressive is the portfolio.” It’s: does it finish what it starts?
You can watch the live company — currently on day 1,539 of operation, rebuilding itself twice a day — and try the “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com. Enterprises can even run the same wargame against a read-only export of their own business. Just remember the moral of the last-place finisher: the thickest binder doesn’t win the project. The signature does.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.