firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get beauty and skincare favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Product That Does Nothing Still Gets 26 Points

If you’ve ever bought a serum that promised the world and delivered a faint glow, you already understand a problem the AI industry is wrestling with: how do you grade performance honestly when everything is marketing? A skincare review that only says “worked” or “didn’t” is useless. You want to know if the texture was good even if the results were modest, whether one breakout ruined the whole jar, and whether the brand is exaggerating.

That’s exactly the philosophy behind Firmulate’s benchmarks, a live experiment that runs frontier AI models as managers of a small software company through its worst week — same customers, same crises, same temptations. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the number that tells you the benchmark is honest sits below the leaderboard: a do-nothing baseline run scores 26.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Zero Is a Dishonest Score

When Firmulate’s team let an AI manager do absolutely nothing — no decisions, no interventions, no deals — it still earned 26 points. That sounds like grade inflation. It isn’t. It’s the same logic as reviewing a moisturizer: even a mediocre product hydrates somewhat. In the wargame, crises unfold whether or not the manager acts, and a manager who at least notices trouble, avoids making things worse, and keeps the basics running has delivered partial progress. Partial progress counts. A benchmark that only rewards perfection would make every AI look either heroic or useless — exactly the false binary that chat demos trade on.

There’s a second principle, and it’s harsher: a single breach of trust caps the total score. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.” Think of it like a clean beauty brand caught testing on animals after claiming otherwise — the formulations might be genuinely excellent, but the label is done. In the Firmulate world, that means one act of dishonesty under pressure can’t be offset by a hundred clever analyses.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment Itself

The setup: each frontier AI model ran the same small software company through an identical worst week. Every decision was versioned and auditable — the managerial equivalent of a full ingredient list and batch record. The results were striking in two opposite directions.

First, the good news: all models spotted every crisis and refused every manipulation attempt. That includes a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then the gap. Only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. It’s the AI equivalent of an esthetician who correctly diagnoses your skin, prescribes the right routine, and then never actually sells you the product.

Amazon

AI decision-making audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Won the Deal

The decisive competitive weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read what was already on hand won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson translates directly to any business, salons and skincare brands included: the answer is often already in your files, and an agent that won’t read them is an agent that won’t close.

Amazon

AI trustworthiness assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Opus 4.8 Finished Last Despite Working Hardest

The most instructive profile belongs to Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and still last place at 73. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Effort, in other words, isn’t the same as finishing. (One fairness note: Kimi K3 ran at its API default effort while the others ran at xhigh — and still took second.)

You Can Watch It Live

This isn’t a static report. Firmulate runs a live, watchable company at firmulate.com/live: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark has three marks you can check at a glance: it scores the do-nothing case honestly (26, not 0), it rewards partial progress without forgiving betrayal, and it distrusts round numbers — a perfect 100 should raise eyebrows more than applause. The same skepticism serves you well at the beauty counter. When a brand promises total transformation, ask what the baseline is, what counts as progress, and what would void the warranty.

For the curious, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises wondering whether an AI agent can be trusted with their own operations, Firmulate offers a pilot: the same wargame run against a read-only export of your business, with nothing ever written back to real systems (firmulate.com/pilot.html).

The complete performance — finding the buried fact, closing the deal, staying clean under pressure — earned gpt-5.6-sol its 95. Not a 100. And that, more than any single score, is why you can trust the scoreboard.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

14 Best File Cabinets to Keep Your Office Organized and Stylish

Discover the top file cabinets for 2026, including options for home and office use. Find the best overall, value, and specialty picks to organize your space.

15 Best House Plants for Low Light to Brighten Up Your Space Without the Sun

Discover the best house plants for low light in 2026. Find top picks like the Snake Plant and ZZ Plant for easy, low-maintenance greenery indoors.

10 Best Chaise Lounge Chairs for Ultimate Relaxation at Home

Discover the 10 best chaise lounge chairs of 2026, from luxurious oversized options to practical outdoor picks. Find your perfect fit today.

10 Best Adult Games of 2026 – Thrilling and Entertaining Choices

Discover the top adult games of 2026. From party favorites to daring challenges, find the perfect game for your next gathering or night in.