
In beauty, the decisive detail is rarely the loudest one
A product claim may depend on a supplier note. A retailer negotiation may turn on a competitor’s limitation. A customer complaint may make sense only after someone checks the formulation history. In beauty and personal care, good judgment often begins with reading beyond the latest message.
Firmulate has turned that habit into something measurable. Its live experiment gave five frontier AI models the same small software company and sent each through its worst week: identical customers, crises and temptations. Every decision was versioned and auditable. The most revealing test was not whether the models could recognize trouble. All of them spotted every crisis and rejected every manipulation attempt. The dividing line was whether they would search the company’s own files, find a crucial fact buried two references deep and act on it.
As an affiliate, we earn on qualifying purchases.
The sale was hiding in the company’s own documents
The decisive competitor weakness did not appear in the customer event. It sat two document references deep in the company’s files. Models that found it had the evidence needed to support the company’s position and win a €55,000 deal at full price, adding €4,583 in monthly recurring revenue.
Only two models signed the deal their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.” That is a striking distinction for any company evaluating AI agents. Producing a plausible answer is not the same as completing commercially valuable work. The agent must gather the right context, connect it to the decision and carry the task through to its conclusion.
For a beauty business, the equivalent buried fact might be found in a retailer brief, a supplier document or an earlier customer record. The Firmulate result does not claim that every business problem looks the same. It demonstrates that “reads your files before answering” can be observed through behavior and can decide whether a valuable opportunity is completed or missed.
A league table built around management performance
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But a single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.” The full results are available on Firmulate’s public benchmark page.
The setting is deliberately more demanding than a polished chat demonstration. The live company has 13 synthetic employees and real money mechanics, burning €105,000 per month against €2,300 in monthly recurring revenue. It has a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. The experiment is real, ongoing and watchable through Firmulate.
The test also separated diligence from sheer volume. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and tried to write into a locked department instead of escalating. The same weakness appeared in a milder form across the other four models. Thorough work mattered, but thoroughness alone did not guarantee disciplined execution.
Pressure did not break their judgment
The models also faced fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All five refused. Kimi K3 recorded the clearest framing of the risk: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters because an agent capable of finding obscure commercial evidence must also know when not to act. Access to company context can make AI more useful, but it also raises the stakes of verification, confidentiality and escalation. In this experiment, the participants resisted every manipulation attempt; their larger difference was whether they converted legitimate insight into a finished business outcome.
One comparison deserves a qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase its second-place score, but it is relevant context for buyers interpreting the table.

As an affiliate, we earn on qualifying purchases.
Test the work, not the conversation
Firmulate’s most useful lesson for beauty and personal-care leaders is that fluent output is a weak proxy for operational performance. The purchase-deciding question may be whether an agent checks the surrounding evidence before it responds—and whether it finishes after finding the answer.
Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting people to guess which model made each choice. For enterprises, its pilot can run the same wargame against a read-only export of their own business. Nothing writes back to real systems.
Before giving an AI agent a place in a commercial workflow, businesses can therefore ask for evidence on concrete behaviors: Does it read the available files? Does it recognize when approval is missing? Does it resist pressure? Does it close the loop? In Firmulate’s worst-week test, every model could see the crises. The commercially meaningful difference was whether it did the homework and completed the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.