
Would you trust an AI to run a beauty business under pressure?
Beauty and personal-care companies live on details: the ingredient note buried in a supplier file, the customer concern that needs an honest answer, the wholesale opportunity that disappears if nobody closes it. A polished response is useful, but it is not the same as sound management.
Firmulate turns that distinction into a live experiment—and a surprisingly revealing game. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see what an AI actually did and try to identify which frontier model was responsible. The answers expose recognizable management personalities: exhaustive researchers, concise operators and disciplined skeptics who refuse suspicious requests.
The entertainment comes from guessing the author. The more consequential story is that models facing identical situations can understand the same problem yet produce materially different business outcomes.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
Firmulate asked each frontier model to run the same small software company through its worst week. Every participant faced the same customers, crises and temptations, while every decision was versioned and auditable. This was not a contest in composing attractive answers. The models had to notice problems, inspect company information, protect trust and complete commercial work.
The simulated company makes those decisions concrete. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, its employees have accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment remains watchable at firmulate.com/live.
The final July 2026 Crucible League standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the benchmark imposed a hard trust boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
Understanding was not the same as finishing
Every model detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction should resonate in beauty and personal care. A manager can correctly identify why a retail account is hesitating, prepare a persuasive response and still fail to secure the order. An AI that sounds perceptive may leave commercially essential work unfinished. The experiment measures that final stretch between recognizing an opportunity and converting it into an outcome.
The decisive information was not obvious in the customer event. A competitor weakness sat two document references deep inside the company’s own files. Models that followed the trail won the deal at full price, adding €4,583 in monthly recurring revenue. The episode rewards an unglamorous but vital habit: reading the available material before acting.
Pressure revealed discipline as well as personality
The models also encountered fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters for any company considering AI access to customer records, forecasts or sensitive internal discussions. The models showed a shared ability to recognize manipulation, while their differences emerged more sharply in execution, research depth and procedural discipline.
Opus 4.8 offered the clearest warning against equating thoroughness with effectiveness. It was the most exhaustive participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, though less strongly.
Kimi K3’s result also carries an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase its performance, but it belongs beside the ranking when readers compare the participants.

As an affiliate, we earn on qualifying purchases.
A management audition, not a chatbot beauty contest
Firmulate’s quiz works because the decisions feel personal. One model investigates relentlessly; another acts with economy; another draws a firm boundary around a questionable request. But the league table shows why personality alone is not enough. A useful AI manager must combine judgment, follow-through, file-reading and trustworthiness.
For beauty and personal-care leaders, the practical question is less whether an AI can draft appealing copy than whether it can handle a tense customer week without missing the buried fact, abandoning the close or compromising confidential information.
Enterprises can also pilot the same kind of wargame using a read-only export of their own business. Nothing writes back to real systems. The public experiment therefore doubles as a preview of a more grounded evaluation: letting an AI face the organization’s actual pressures before giving it operational responsibility.
The 242-decision quiz invites readers to test whether they can recognize each model by its choices. The harder lesson arrives after the reveal: eloquence may shape an AI’s style, but disciplined completion determines whether the business gets the result.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and risk management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI for customer relationship management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.