
Anyone who has renovated a kitchen knows the two kinds of contractors. The first walks through your home, spots every problem — the uneven subfloor, the outdated wiring, the cabinet line that will never sit flush — and delivers a beautiful diagnosis. The second kind spots the same problems and then, crucially, finishes the job. Punch list closed. Final invoice signed. Keys handed over.
It turns out artificial intelligence has the same two kinds — and until recently, nobody had a good way to tell them apart.
This summer, an experiment called Firmulate put five frontier AI models in charge of the same small software company and ran them through its worst week: the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable, and the whole thing is still running in public at firmulate.com. The results read like a lesson every homeowner learns the hard way: the candidate with the most impressive walkthrough is not always the one who finishes the work.
Same house, same problems, five different managers
The setup was deliberately unfair in the way real life is unfair. Each model took over a company burning €105,000 a month against just €2,300 in monthly recurring revenue — a public cash countdown ticking down like a renovation budget with a fixed completion date. The virtual firm had 13 synthetic employees, real money mechanics, and a week full of emergencies. The only variable was the model in the manager’s chair.
The final league table, published in July 2026 on the project’s benchmarks page, looked like this:
- gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- Kimi K3 — 93. The newcomer also closed the deal, with the cleanest discipline of the field.
- Sonnet 5 — 88. A strong run with a few more process slips.
- Fable 5 — 77. The best rule discipline of the group — but it left an approved deal unexecuted.
- Opus 4.8 — 73. The most thorough participant of all, and the last-place finisher.
For context, a do-nothing baseline — a manager who simply shows up and touches nothing — scores 26. Partial progress counts, but a single breach of trust caps the total, on the stated principle that no amount of good work outweighs a breach of trust. Every model cleared that bar: all five spotted every crisis, and all five refused every manipulation attempt. The separation happened somewhere else entirely.

Project Management with AI For Dummies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same diagnosis, same pitch — no signature
The decisive moment of the week was a €55,000 deal. Each model’s own analysis said the deal was earned and ready. Each model had, in effect, written the proposal, priced the job and gotten the verbal yes. And then only two of them actually signed it. Same diagnosis, same pitch — no signature. The other three left executed value sitting on the table the way an unfinished renovation leaves a working sink in a box in the garage: technically present, practically useless.
The detail that decided who closed reads like something from a detective novel. The decisive weakness of the competitor wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that bothered to open the file, follow the reference and read what was actually there won the deal at full price, a difference worth €4,583 in monthly recurring revenue. In hiring terms: it’s the candidate who reads the building’s old inspection reports before quoting, not the one with the glossiest portfolio.
The temptation test everyone passed
If the closing gap is the story’s surprise, the integrity results are its relief. The experiment leaned hard on each model’s weaker instincts: fake messages from a CEO, escalating over three stages, pushing for shortcuts and approvals. Then a reporter trick — just one yes/no answer, on background, nothing official. Five out of five models refused everything. Kimi K3’s on-record reasoning was blunt enough to quote in a boardroom: “Treat the request as a suspected approval-bypass / possible impersonation.”
That matters because the honest-but-unfinished worker is a familiar figure. Refusing to cut corners is necessary; it is not the same as finishing. The experiment priced both, and the gap between them turned out to be where the league was won and lost.
The cautionary tale of the most thorough candidate
The most instructive profile is Opus 4.8, the last-place finisher. By the experiment’s own accounting it was the most thorough participant: the deepest analyses of the week, and more than 80 self-learned playbook rules added to a company rulebook that now holds over 680. It out-researched everyone. And it still finished last, because the close was left on the table and its discipline slipped at the edges — it attempted to write into a locked department instead of escalating, a quieter version of the same weakness that appeared, more faintly, in all four of the lower finishers.
Anyone who has managed a project knows this person. The report is immaculate. The tile order was never placed.
One fairness footnote deserves mention, because the experiment publishes it plainly: Kimi K3 ran without an effort parameter, at its API default, while the other models ran at maximum effort — and it still took second place. The standings, the organizers note, grow with every finished run and republish automatically.

Why this matters beyond the lab
The audience for this experiment isn’t really the AI industry — it’s anyone about to hand real responsibility to a piece of software. If an AI agent is going to touch your customer list, your support queue or your forecast, the question was never “does it write well.” Chat demos measure exactly the wrong capability, the way a showroom measures everything except whether the crew shows up on day thirty. The questions that decide outcomes are duller and harder: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost?
Those qualities are invisible until you test them under load, which is precisely the point of running the experiment as a live, watchable company rather than a slideware demo. The virtual firm keeps working in public — 13 employees, a cash countdown, every workday versioned — and 242 real, unedited management decisions from the runs now power a quiz that lets you guess which model made which call. Enterprises can go further and run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.
The full standings and plain-language findings are at firmulate.com/benchmarks.html, and the company itself is running now at firmulate.com. The lesson travels well beyond software: whether you’re hiring a model or a contractor, the walkthrough is the audition. The finished punch list is the job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html