firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Beauty and personal care companies run on details that are easy to miss: a customer ready to leave, a sales opportunity buried in internal files, or a message that looks urgent but should not be trusted. Firmulate’s latest experiment put AI models in charge of a small software company facing that kind of pressure. The result suggests that polished answers are only part of the job: an AI workforce also has to read carefully, make sound decisions and follow through.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get beauty and skincare favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate ran five frontier models through the same difficult week at a software company, with the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workdays are versioned and auditable, and its playbook contains more than 680 self-learned rules.

The experiment is live and watchable at Firmulate. The point is to examine management quality in action: what models do when customer relationships, company finances and trust are on the line.

Amazon

AI customer service chatbot for beauty brands

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The newcomer takes second place

In the final July 2026 Crucible League, Moonshot’s Kimi K3 scored 93, placing second behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The benchmark counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That difference matters: recognizing the right move is not the same as completing it.

Amazon

beauty brand sales analytics tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The crucial clue was buried

The deal hinged on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found the clue, closed the deal and saved the churning customer. It also resisted all three baits and made only one deviation, giving it the cleanest discipline in the field.

The pressure included fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

customer relationship management software for beauty industry

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness does not guarantee follow-through

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still finished last. The deal was left unsigned, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four other models.

That gap between diagnosis and execution is relevant well beyond software. For beauty and personal care businesses considering AI in customer service, sales or forecasting, a persuasive response is not enough. The model has to consult the information available, protect trust and carry a sound decision through to completion.

Amazon

AI-driven sales forecasting tools for personal care

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A result worth testing against your own work

The league table shows how close the top two results were: K3 finished two points behind gpt-5.6-sol, while placing ahead of three of the four Western models in the field. That makes model choice a practical question for businesses, not a decision to settle by reputation alone. Firmulate offers the full results and plain-language findings at its benchmark page.

Readers can also try a “guess the model” quiz built from 242 real, unedited management decisions at Firmulate. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

What business leaders can take away

Firmulate’s experiment shows why AI evaluations should look at decisions and follow-through, not just chat quality. K3’s second-place result, buried-file discovery and clean conduct make a strong case for testing models on the work your business actually needs them to do.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

13 Best Standing Fans to Keep You Cool All Summer Long

Discover the top standing fans of 2026. Find the best overall, value, premium options, and more to keep your space cool all year round.

Gildan Activewear Surges In Global Coverage

Gildan Activewear’s media mentions have increased sharply, signaling heightened global interest in the company’s activities and market position.

The Next AI Benchmark Belongs in the Executive Suite

AI agents may ace chat tests yet fail the management test: reading buried evidence, closing earned deals, and staying honest under sustained pressure.

StrongMocha News Group Expands into Nanotechnology

AIThis post was created with the assistance of artificial intelligence (AI).Berlin, Germany…