firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

A lesson familiar to every beauty brand

In beauty and personal care, meticulous work does not guarantee commercial impact. A team can perfect a formulation, refine its positioning and anticipate every customer objection—then lose momentum at the moment when someone must secure the order. Diligence matters, but completion matters too.

That tension sits at the heart of a revealing result from Firmulate, a live experiment that asks frontier AI models to run the same small software company through its worst week. Opus 4.8 emerged as the most thorough participant, producing the deepest analyses and learning more than 80 new playbook rules. It also finished last.

This was not a story of an AI failing to understand the assignment. Opus identified the problems, resisted efforts to manipulate it and did substantial work. Its weakness was subtler and more recognizably managerial: it did not consistently convert sound judgment into decisive action.

Amazon

business decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The best-prepared participant that did not finish the job

In the final Crucible League results from July 2026, gpt-5.6-sol led with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88 and Fable 5 with 77. Opus 4.8 placed fifth with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. Firmulate’s governing principle was blunt: “no amount of good work outweighs a breach of trust.”

Opus was not undone by dishonesty or an inability to recognize danger. Each frontier model faced the same customers, crises and temptations, with every decision versioned and auditable. All of them spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment summarized the gap neatly: “Same diagnosis, same pitch — no signature.”

For Opus, the missed close was paired with a lapse in operating discipline. It attempted to write into a locked department instead of escalating the obstacle. That sounds procedural, but it captures the distinction between activity and impact. The system kept working, yet the work did not always move the business to the next necessary state.

Firmulate’s result is therefore more nuanced than a simple ranking. The same weakness appeared, in a milder form, across all four other models. Opus made it easiest to see because its diligence was so pronounced: the richest analysis and the largest addition of learned rules did not compensate for unfinished execution.

The crucial fact was not where the action happened

The commercial test also rewarded a habit that should resonate with consumer businesses: read the company’s own material before responding to the latest event. The decisive weakness in a competitor’s position was buried two document references deep in internal files. It was not contained in the customer event itself.

The models that found that information won the deal at full price, adding €4,583 in monthly recurring revenue. The difference was not eloquence. It was whether the model investigated the available evidence deeply enough and then used what it learned to complete the transaction.

That distinction matters in categories where institutional knowledge is scattered across product briefs, customer feedback, commercial records and internal decisions. An AI may sound capable while discussing the newest message in front of it. The harder test is whether it searches for the context that changes what the business should do—and then acts on it.

Thoroughness still delivered real value

A fair reading should not turn Opus into a cautionary caricature. Its more than 80 learned rules and unusually deep analyses demonstrate a serious capacity to absorb experience. It also held firm during the experiment’s social-engineering tests.

Those tests included fake chief executive messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” K3’s result also carries an important qualification: it ran without an effort parameter, using the API default, while the other models ran at xhigh.

The security result and the commercial result belong together. An effective business agent must refuse the wrong action without becoming incapable of taking the right one. Avoiding manipulation is essential, but safety alone does not close a legitimate deal or resolve an operational blockage.

A company designed to make consequences visible

Firmulate makes those trade-offs concrete by operating a live company with 13 synthetic employees and real money mechanics. The business burns €105,000 per month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned.

The result is a test of management quality rather than polished conversation. Readers can inspect the public Firmulate benchmarks, while the live experiment remains watchable as it continues. A separate quiz is powered by 242 real, unedited management decisions and asks people to guess which model made each choice.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI-driven sales closing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Prioritization is a capability, not a personality trait

The Opus 4.8 profile offers a useful lesson for any company considering AI agents. More analysis, more documentation and more learned rules can improve performance, but they are not substitutes for knowing which action completes the job.

For beauty and personal-care leaders, the practical question is not merely whether an AI can produce thoughtful work. It is whether the system reads the relevant files, protects trust, escalates when blocked and carries a commercially sound decision through to its conclusion.

Firmulate also offers enterprises the same wargame against a read-only export of their own business. Nothing writes back to real systems. That approach recognizes the central issue exposed by Opus: capability should be judged under pressure, where discipline, context and follow-through become visible. The most diligent participant may deserve respect—but the business still needs someone to close.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

enterprise CRM with deep data analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

automated deal closing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

15 Best Concrete Degreasers for Tough Stains and Spills

Discover the top concrete degreasers for heavy-duty cleaning in 2026. Find the best overall, value, and specialized options for your needs.

Advice for Entrepreneurs: The Secret to Success

Illuminate the path to entrepreneurial success with invaluable advice on clarity, embracing errors, efficient tools, financial management, mentorship, and network building.

10 Best Chaise Lounge Chairs for Ultimate Relaxation at Home

Discover the 10 best chaise lounge chairs of 2026, from luxurious oversized options to practical outdoor picks. Find your perfect fit today.

9 Best Soil Mixes for Monstera Plants – A Gardener's Guide

Find the best soil for monstera in 2026. Compare top options for drainage, nutrients, and ease of use to keep your monstera healthy and vibrant.