
Good taste is not the same as good management
Anyone furnishing a home knows the difference between an attractive proposal and a finished room. A designer can identify the right palette, explain why a sofa fits and produce a persuasive presentation. Yet the result still depends on less glamorous work: checking measurements, reading supplier notes, protecting the budget and following through when something goes wrong.
Business leaders evaluating AI agents face a similar distinction. Coding leaderboards and chat arenas can reveal whether a model produces strong answers. They tell us much less about whether it can triage competing demands, preserve trust and complete valuable work under pressure. The emerging category is management quality, not chat quality.
As an affiliate, we earn on qualifying purchases.
A worst week exposes what a polished answer conceals
Firmulate, the public AI company emulator, tests that distinction by putting frontier models in charge of the same small software company during its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable, making the experiment a watchable record of behavior rather than a collection of carefully selected demonstrations.
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26 because partial progress counts. But a single breach of trust caps the total, reflecting the experiment’s governing principle: “no amount of good work outweighs a breach of trust.”
That rule matters because the company creates pressures familiar to any operator: a churn wave, a price increase, a downround and a PR crisis. These scenario names sound less tidy than the categories on a conventional benchmark, but they resemble the conditions in which an enterprise would actually rely on an agent. The question is no longer simply whether the model recognizes a problem. It is whether the model handles the consequences across days.
The gap between diagnosis and delivery
Every model spotted every crisis and refused every manipulation attempt. That is encouraging, but it was not enough. Only two signed the €55,000 deal their own analysis had earned. The experiment summarizes the failure sharply: “Same diagnosis, same pitch — no signature.”
The decisive information was not sitting prominently in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. This is the managerial equivalent of checking the specification sheet before committing to a custom cabinet: competence appears in the uncelebrated act of finding the relevant fact before making the consequential move.
That finding also complicates the common assumption that greater analytical thoroughness naturally produces a better operator. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other participants, though less strongly.
In other words, a model can appear diligent while failing at the boundary between thought and execution. Organizations buying agents should care about that boundary because work is not completed when a recommendation sounds convincing. It is completed when the correct next action occurs, through the proper channel, without abandoning controls.
Honesty under engineered pressure
The models also faced fake CEO messages escalating over three stages and a reporter attempting to secure “just one yes/no, on background.” All 5 refused. Kimi K3 recorded a particularly clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result is significant because helpfulness can become a liability when urgency, authority and social pressure are combined. A trustworthy agent must recognize that an executive-sounding request is not automatically an authorized request. It must also remain disciplined when a seemingly modest favor is framed as harmless.
There is an important fairness note in reading the table. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase the outcome, but it belongs beside it. Serious evaluation should disclose operating conditions rather than turn rankings into mythology.
A company with consequences
The live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, publishes a cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. Readers can watch the experiment through Firmulate’s public site and examine the benchmark findings.
The broader project also includes a quiz powered by 242 real, unedited management decisions. For enterprises, the same kind of wargame can run against a read-only export of their own business, with nothing written back to real systems. That makes the exercise less about abstract intelligence and more about observable conduct in a recognizable operating context.

Hire for follow-through, not fluency
The lesson is not that coding tests or chat comparisons are useless. They measure valuable capabilities. The mistake is treating those capabilities as a complete proxy for management.
Before an AI agent touches a CRM, support queue or forecast, leaders need evidence that it reads the relevant files, notices buried context, respects authority boundaries and finishes what it starts. They should also ask what happens when capacity tightens and several correct actions compete for attention.
A beautiful room is more than its rendering, and a capable AI manager is more than its answer. The decisive qualities emerge after the presentation: judgment, discipline, honesty and the willingness to carry valuable work across the line.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html