
Beauty and personal care companies run on details that are easy to miss: a customer ready to leave, a sales opportunity buried in internal files, or a message that looks urgent but should not be trusted. Firmulate’s latest experiment put AI models in charge of a small software company facing that kind of pressure. The result suggests that polished answers are only part of the job: an AI workforce also has to read carefully, make sound decisions and follow through.
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate ran five frontier models through the same difficult week at a software company, with the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workdays are versioned and auditable, and its playbook contains more than 680 self-learned rules.
The experiment is live and watchable at Firmulate. The point is to examine management quality in action: what models do when customer relationships, company finances and trust are on the line.
AI customer service chatbot for beauty brands
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The newcomer takes second place
In the final July 2026 Crucible League, Moonshot’s Kimi K3 scored 93, placing second behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The benchmark counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That difference matters: recognizing the right move is not the same as completing it.
beauty brand sales analytics tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The crucial clue was buried
The deal hinged on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found the clue, closed the deal and saved the churning customer. It also resisted all three baits and made only one deviation, giving it the cleanest discipline in the field.
The pressure included fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
customer relationship management software for beauty industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness does not guarantee follow-through
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still finished last. The deal was left unsigned, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four other models.
That gap between diagnosis and execution is relevant well beyond software. For beauty and personal care businesses considering AI in customer service, sales or forecasting, a persuasive response is not enough. The model has to consult the information available, protect trust and carry a sound decision through to completion.
AI-driven sales forecasting tools for personal care
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A result worth testing against your own work
The league table shows how close the top two results were: K3 finished two points behind gpt-5.6-sol, while placing ahead of three of the four Western models in the field. That makes model choice a practical question for businesses, not a decision to settle by reputation alone. Firmulate offers the full results and plain-language findings at its benchmark page.
Readers can also try a “guess the model” quiz built from 242 real, unedited management decisions at Firmulate. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

What business leaders can take away
Firmulate’s experiment shows why AI evaluations should look at decisions and follow-through, not just chat quality. K3’s second-place result, buried-file discovery and clean conduct make a strong case for testing models on the work your business actually needs them to do.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
