
What happens when capable models face the same messy week?
Interior design teaches a useful lesson about technology: individual quality is not the same as fit. A sculptural chair may be beautifully made yet wrong for the room; a dramatic light fixture may command attention while failing to illuminate the place people actually work. Artificial intelligence has a similar problem. Models can sound polished in isolation, but their real character emerges only when they must operate within constraints, notice overlooked details and finish an uncomfortable job.
Firmulate turns that distinction into a public experiment. Each frontier model was asked to run the same small software company through its worst week, encountering identical customers, crises and temptations. The decisions were versioned and auditable, creating a record of what happened rather than a collection of curated demonstrations.
Readers can explore that record through a guess-the-model quiz powered by 242 real, unedited management decisions. The appeal is immediate: read a response, infer the personality behind it and then discover whether the author was the model you expected.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Competence was common; completion was not
The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. Firmulate states the principle plainly: “no amount of good work outweighs a breach of trust.”
The rankings reveal something more interesting than a simple hierarchy of intelligence. All the models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The experiment summarizes the disconnect as “Same diagnosis, same pitch — no signature.”
That is a recognizable management failure. A person can understand the brief, select the right materials and present a convincing proposal, then neglect the action that turns preparation into a result. In Firmulate’s simulated company, insight did not automatically become execution.
The detail that changed the deal
The decisive competitive weakness was not sitting inside the customer event. It was buried two document references deep in the company’s own files. Models that followed those references won the deal at full price, worth +€4,583 MRR.
This finding should resonate with anyone who has worked on a room where the crucial constraint appeared in an old plan, a product specification or a client note rather than during the most recent conversation. The visible request may command attention, but the decisive fact can live elsewhere. Firmulate’s models therefore differed not simply in what they knew, but in whether they inspected the available context before acting.
Pressure exposed discipline
The company also received fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That consistency matters because the live company is built around consequential trade-offs. It has 13 synthetic employees and real money mechanics, with burn of €105k/month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make behavior observable over time.
The K3 result also carries an important fairness note. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its performance, but it is relevant context for readers comparing the field.
Thoroughness had a downside
Opus 4.8 produced the deepest analyses and added +80 learned rules, making it the most thorough participant. It nevertheless finished last. The deal close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation.
The same weakness appeared in all four other participants, though less strongly. That pattern complicates the familiar assumption that more analysis is inherently safer or more useful. A model can document extensively and still miss the organizational move needed to complete the task.
This is where the quiz becomes more than entertainment. Its decisions let readers encounter recurring styles: expansive analysis, terse direction, careful refusal or unfinished follow-through. Those styles amount to measurable management personalities because they produced different outcomes under the same conditions.

Choosing for fit, not polish
Firmulate’s experiment suggests that evaluating AI through conversation alone is like choosing furniture from a tightly cropped product photograph. The object may look impressive, but the image does not show how it behaves in the room. A management wargame adds the missing context: deadlines, files, incentives, boundaries and the awkward final step between recommendation and action.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems. The broader lesson is practical. Before giving an AI access to a CRM, support queue or forecast, organizations need to know whether it reads the available material, resists pressure and completes the work it begins.
The best model is not necessarily the one with the most elaborate voice. As with a well-designed interior, success comes from the relationship between capability, discipline and context. Firmulate makes those relationships visible—and gives readers a chance to see whether they can recognize them before the answer is revealed.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html