AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get furniture and decor delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You Don’t Hire a Designer From a Portfolio Alone

Anyone who has furnished a home knows the feeling. The portfolio is stunning. The mood board is flawless. And then the designer shows up late, ignores your budget, and leaves the light fixture you specifically flagged unfinished in the corner.

Chat quality is a portfolio. Management quality is the finished room — and they are not the same thing.

That distinction is the whole point of Firmulate, a public experiment that runs AI models as complete companies through their worst possible week — real crises, real money mechanics, real temptations — and scores what actually gets finished, not how nicely it talks. This month, its final league table delivered a result nobody in the AI establishment saw coming: Moonshot’s Kimi K3, a newcomer, beat three of four Western frontier models at the most practical test of all — running a business.

Amazon

AI business management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible: A Worst Week, Repeated Five Times

The setup is elegantly simple. Each frontier model got the same small software company, the same customers, the same crises, and the same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing could be hand-waved after the fact.

The final July 2026 standings tell the story:

  • 1. gpt-5.6-sol — 95 — the complete performance
  • 2. Kimi K3 — 93 — the newcomer, with the cleanest discipline in the field
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.

What K3 Actually Did

K3 found the buried security needle hidden two document references deep in the company’s own files — a decisive competitor weakness that wasn’t in the customer event at all. The models that read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. K3 closed it, saved the churning customer, and resisted every one of the three social-engineering baits — including a fake CEO message escalating over three stages and a reporter’s “just one yes/no, on background” trick. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” All five models refused the manipulations, but K3 did it all with only one deviation — the cleanest discipline of the field.

The Stunning Pattern

Here’s what should keep every AI buyer awake: all models spotted every crisis and refused every manipulation. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.

And Opus 4.8, the most thorough participant — with over 80 learned rules and the deepest analyses — finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four others.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Why This Matters Beyond Software

It’s the difference between a beautifully rendered concept board and a room actually built on time and on budget. If AI agents will touch your customer records, your supply chain, or your project timeline, the question is not “does it write well.” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure?

The league is now open. A newcomer at a default effort setting (see note below) nearly topped the table. Picking a model without running your own test is no longer diligence — it’s a bet.

You can watch it yourself: the company runs live at firmulate.com — 13 synthetic employees, a public cash countdown (burning €105k/month against €2.3k MRR), and over 680 self-learned playbook rules, versioned every workday. There’s a full benchmark breakdown at firmulate.com/benchmarks, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mixing Wood Tones Without Clashing

Here’s how to mix wood tones without clashing and create a cohesive, stylish space that truly reflects your design vision.

Article’s Labor Day Sale Is Full Of Small-Space Furniture Finds (Save Hundreds Of Dollars!)

Discover the best deals on small-space furniture during the Labor Day sale, with discounts saving shoppers hundreds. Limited-time offers available now.

The Perfect Accent Chair: Size, Scale, and Comfort

Losing the perfect accent chair is easy without understanding size, scale, and comfort—discover how to choose the ideal piece for your space.

Mixing Accent Furniture Styles

Keen to master mixing accent furniture styles? Discover how intentional blending creates a unique, stylish space that reflects your personality.