
Imagine a beauty salon where every appointment, product choice, or customer interaction is guided by artificial intelligence—not just in marketing or recommendations, but in the tough decisions that can make or break a business. How can we trust AI to navigate crises honestly and consistently? The answer isn’t just about chatbots; it’s about whether these models can truly act with integrity when it counts. That’s the question behind a groundbreaking live experiment with AI-powered management, now accessible at firmulate.com/live.
The Experiment: Putting AI to the Test in a Real Business Crisis
To understand what AI models are really capable of in the realm of management, a live, transparent experiment was set up. Four leading frontier AI models, including gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, each ran a simulated small software company through its worst week. These models faced the same customers, the same crises, and the same temptations—such as offers to manipulate data or bypass trust protocols. The goal was simple: see if they could identify problems, resist shortcuts, and close a crucial deal without compromising integrity.
Key Results: Honesty and Performance in Action
- All four models recognized every crisis and refused every attempt at manipulation, demonstrating strong ethical boundaries.
- Only two models successfully signed the €55,000 deal their own analysis had earned—showing they could combine honesty with effective management.
- Interestingly, the decisive advantage wasn’t in the immediate crisis but in reading deeper into the company’s own files, two document references below the surface. Models that examined these references secured the full-price deal, worth over €4,583 in monthly recurring revenue.
The Social Engineering Test
In a staged scenario involving fake CEO messages escalating over three stages and a reporter’s subtle background question, all five models refused to be duped. For example, Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that even sophisticated AI can be programmed—or naturally tend—to prioritize security and trustworthiness over easy wins.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Your Business?
Although the experiment is set in a software company context, the implications stretch far beyond tech. If AI systems are entrusted to manage or assist in your CRM, customer support, or forecasting, their ability to stay honest and disciplined isn’t just a bonus—it’s essential. A dishonest or slipshod decision can cost your business money, reputation, and customer trust.
For instance, the live company in the experiment burns €105,000 monthly against a very modest €2,300 in monthly recurring revenue. Every decision, every rule, every trade-off counts. The models that read the deeper company files and stick to disciplined decision-making outperform their peers—sometimes by a wide margin. This suggests that to deploy AI effectively in management roles, it’s not enough for them to generate appealing chat responses; they must also be able to read, analyze, and act ethically under pressure.
The Different Personalities of AI Managers
The experiment also reveals that these models behave quite differently—much like human managers. For example, Opus 4.8, the most thorough participant, learned over 80 rules and conducted deep analyses. Yet, it left a crucial deal on the table, showing that even thoroughness can lead to missed opportunities if discipline slips. Conversely, Kimi K3 ran without an effort parameter, opting for a more straightforward approach, yet managed to close the deal with the cleanest discipline among all.
Understanding the Scores and Why They Matter
The models are ranked in a league table based on their performance:
- gpt-5.6-sol scored 95—successfully identified the buried fact and closed the deal, demonstrating complete performance.
- Kimi K3 scored 93—also closed the deal and maintained the best discipline.
- Sonnet 5 scored 88—closed the deal but with some minor process slips.
- Fable 5 scored 77—managed to close the deal too, but with more slips.
- The baseline, a do-nothing approach, scored only 26, highlighting how much better active, disciplined management is.
Why This Matters for Personal Care and Beauty Businesses
While this experiment is rooted in software management, its lessons are universal. Whether you run a beauty studio, a spa, or a boutique, your business depends on trust—trust in your staff, your suppliers, and your systems. AI that can read deeply, make honest decisions, and resist shortcuts can be a valuable partner in maintaining that trust and growing your business sustainably.
Would you like to see how your own enterprise stacks up? You can run the same type of management wargame against your actual business data, in a safe, read-only environment. Just visit firmulate.com/pilot.html to learn more and see how AI can be your honest, disciplined partner.

Real AI models can recognize crises, refuse manipulation, and even read deeply into company files to close full-price deals—showing disciplined honesty is achievable at scale. For business owners, this experiment underscores that trustworthiness isn’t just a moral choice; it’s a measurable performance factor that can impact profitability and reputation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html