
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Product That Does Nothing Still Gets 26 Points
If you’ve ever bought a serum that promised the world and delivered a faint glow, you already understand a problem the AI industry is wrestling with: how do you grade performance honestly when everything is marketing? A skincare review that only says “worked” or “didn’t” is useless. You want to know if the texture was good even if the results were modest, whether one breakout ruined the whole jar, and whether the brand is exaggerating.
That’s exactly the philosophy behind Firmulate’s benchmarks, a live experiment that runs frontier AI models as managers of a small software company through its worst week — same customers, same crises, same temptations. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the number that tells you the benchmark is honest sits below the leaderboard: a do-nothing baseline run scores 26.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Zero Is a Dishonest Score
When Firmulate’s team let an AI manager do absolutely nothing — no decisions, no interventions, no deals — it still earned 26 points. That sounds like grade inflation. It isn’t. It’s the same logic as reviewing a moisturizer: even a mediocre product hydrates somewhat. In the wargame, crises unfold whether or not the manager acts, and a manager who at least notices trouble, avoids making things worse, and keeps the basics running has delivered partial progress. Partial progress counts. A benchmark that only rewards perfection would make every AI look either heroic or useless — exactly the false binary that chat demos trade on.
There’s a second principle, and it’s harsher: a single breach of trust caps the total score. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.” Think of it like a clean beauty brand caught testing on animals after claiming otherwise — the formulations might be genuinely excellent, but the label is done. In the Firmulate world, that means one act of dishonesty under pressure can’t be offset by a hundred clever analyses.
As an affiliate, we earn on qualifying purchases.
The Experiment Itself
The setup: each frontier AI model ran the same small software company through an identical worst week. Every decision was versioned and auditable — the managerial equivalent of a full ingredient list and batch record. The results were striking in two opposite directions.
First, the good news: all models spotted every crisis and refused every manipulation attempt. That includes a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then the gap. Only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. It’s the AI equivalent of an esthetician who correctly diagnoses your skin, prescribes the right routine, and then never actually sells you the product.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Won the Deal
The decisive competitive weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read what was already on hand won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson translates directly to any business, salons and skincare brands included: the answer is often already in your files, and an agent that won’t read them is an agent that won’t close.
As an affiliate, we earn on qualifying purchases.
Why Opus 4.8 Finished Last Despite Working Hardest
The most instructive profile belongs to Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and still last place at 73. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Effort, in other words, isn’t the same as finishing. (One fairness note: Kimi K3 ran at its API default effort while the others ran at xhigh — and still took second.)
You Can Watch It Live
This isn’t a static report. Firmulate runs a live, watchable company at firmulate.com/live: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day.

The Takeaway
An honest benchmark has three marks you can check at a glance: it scores the do-nothing case honestly (26, not 0), it rewards partial progress without forgiving betrayal, and it distrusts round numbers — a perfect 100 should raise eyebrows more than applause. The same skepticism serves you well at the beauty counter. When a brand promises total transformation, ask what the baseline is, what counts as progress, and what would void the warranty.
For the curious, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises wondering whether an AI agent can be trusted with their own operations, Firmulate offers a pilot: the same wargame run against a read-only export of your business, with nothing ever written back to real systems (firmulate.com/pilot.html).
The complete performance — finding the buried fact, closing the deal, staying clean under pressure — earned gpt-5.6-sol its 95. Not a 100. And that, more than any single score, is why you can trust the scoreboard.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
