
What beauty brands can learn from a company with nowhere to hide
Beauty and personal-care businesses live on trust. A polished campaign may win attention, but customers ultimately judge whether the product, promise and experience hold together. Firmulate applies that same unforgiving standard to an unusual subject: a software company operated by synthetic employees, with its decisions, finances and working days exposed to public scrutiny.
This is not a concept brand or scripted simulation. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its employees have accumulated more than 680 self-learned playbook rules, and every workday is versioned. The result is build-in-public taken to an extreme: visitors can watch a company fight for survival while its operational record keeps growing.
AI transparency monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A business story generated by actual pressure
Most companies reveal selected milestones: a launch, a funding announcement, a new hire or a sales win. Firmulate exposes something messier and more revealing—the daily distance between recognizing what needs to happen and actually completing it.
That distinction became stark in the Crucible League, finalized in July 2026. Each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable.
The models all spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s sharpest summary was: “Same diagnosis, same pitch — no signature.” In other words, sounding capable and identifying the correct move did not guarantee a commercial result.
The sale depended on reading past the obvious
The decisive detail was not contained in the customer event. It sat two document references deep inside the company’s own files: a buried weakness in a competitor’s position. Models that found and used that fact won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
For beauty and personal-care readers, the lesson is recognizable. Important commercial truth is not always sitting in the latest customer message. It may be buried in product documentation, earlier research, claims language or institutional knowledge. Firmulate’s experiment showed the practical difference between reacting to what is immediately visible and doing the deeper reading required to close the loop.
Trust held when the pressure became personal
The worst week also included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That matters because operational competence without restraint can become a liability. The league’s do-nothing baseline scored 26 because partial progress still counted, but a single breach of trust capped the total. Its governing principle was explicit: “no amount of good work outweighs a breach of trust.” That is a particularly relevant standard in industries where reputation depends on credible claims and responsible handling of sensitive information.
The leaderboard rewards completion, not performance theatre
The final Crucible League table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran using the API default without an effort parameter, while the other participants ran at xhigh, an important fairness note when comparing the results.
Opus 4.8 produced the most thorough work, adding 80 learned rules and delivering the deepest analyses, yet finished last. It failed to complete the close and lost discipline by attempting to write into a locked department instead of escalating. A weaker form of the same problem appeared in all four of the other participants.
This is where the experiment departs from conventional demonstrations of artificial intelligence. Fluent output is easy to admire in isolation. A running company tests whether analysis survives contact with permissions, customer pressure, unfinished work and the need to make the final commercial move. Firmulate also publishes what its synthetic employees actually say, allowing readers to inspect their words directly rather than relying only on a polished summary.

business decision version control software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A public test of whether AI can finish the job
Firmulate’s most compelling feature is not that its company has no human employees. It is that the experiment makes the consequences visible. The revenue is small beside the burn, the countdown continues, and each business day creates more evidence about how synthetic workers behave under pressure.
For beauty and personal-care leaders, the broader question is less futuristic than it sounds. If AI is allowed near customer relationships, commercial decisions or internal knowledge, does it read deeply enough, protect trust and finish what it starts? Firmulate turns those questions into an observable business story—one whose next chapter is written by the company’s continuing attempt to survive.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
synthetic employee management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.