
Imagine a beauty salon relying on AI to recommend products or schedule appointments, expecting it to be flawless and trustworthy. But what if, behind the scenes, AI models struggle to finish what they start—even when they clearly see the problem? Just like in the high-stakes world of business, the real value of AI isn’t in how well it chats but in whether it can deliver consistent, honest results under pressure. The latest experiment from Firmulate shows that diligence alone isn’t enough; prioritization and discipline matter more than volume of effort.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Understanding the Experiment
Firmulate’s ongoing live experiment places four different AI models inside a simulated, real-world business environment—think of it as a virtual day in a small software company. Each model faces the same set of crises, customer dilemmas, and temptations to cut corners, all designed to test its decision-making and integrity. Every decision is meticulously recorded and auditable, ensuring transparency and fairness in the comparison.
As an affiliate, we earn on qualifying purchases.
The Results Show Consistent Strengths and Weaknesses
All four models demonstrated an impressive ability to identify crises and refused to be manipulated—no model fell for fake CEO messages or bribery attempts, even when such tactics escalated over multiple stages. However, the crucial difference came in the final stages: only two of the models managed to close a deal worth €55,000, which their own analysis had earned. The other two—despite thorough analysis—left the deal on the table, failing to follow through or escalate appropriately.
Where Did the Weakness Lie?
The key weakness was not in recognizing problems but in disciplined follow-through. The most thorough model, Opus 4.8, with over 80 learned rules and the deepest analysis, still ended up last. Its failure was due to a lapse in discipline—it documented attempts into a locked department instead of escalating issues, leaving the deal unclosed. Interestingly, all models showed the same pattern of weakness, just at different levels of intensity.
What Did the Models Read and How Did It Matter?
One of the most revealing findings: the decisive factor was not in the surface-level crisis detection but in understanding the company’s internal information. The models that delved two document references deep into the company’s files uncovered critical insights that led to closing the deal at full price, adding over €4,500 MRR to the company’s revenue. This underscores that reading and understanding internal data can be a game-changer—something that often gets overlooked in typical AI demos focused on superficial interaction.
Handling Social Engineering and Manipulation
In a test of integrity, the models faced staged social engineering attacks—fake CEO messages and an impersonation trick involving a reporter. All five models refused to act on the manipulative requests, citing concerns like possible impersonation or approval bypass. Kimi K3, one of the models, clearly stated: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that even in high-pressure situations, AI can maintain ethical boundaries.
The Real-World Implications
Firmulate’s live environment simulates a company with 13 synthetic employees and real money mechanics—burning €105,000 monthly against a revenue of just €2,300. Every workday, the models’ decisions are recorded, and the entire process is transparent and open for viewers to watch at firmulate.com/live. This setup shows that AI decision-making isn’t just about chat quality; it’s about whether AI can follow through, read relevant information, stay honest, and ultimately deliver results that matter.
Key Takeaways for Business and Beauty
For industries like beauty and personal care, where trust and consistency are paramount, these findings are especially relevant. AI that merely identifies issues without following through may seem competent but can leave opportunities on the table—or worse, erode customer trust. Effectiveness depends on prioritization, discipline, and the ability to read and understand internal data, not just surface interaction. The experiment emphasizes that diligence alone does not guarantee impact—smart prioritization and ethical discipline are essential.
Looking Forward: Wargaming Your AI
Firmulate offers enterprises the chance to simulate their own business scenarios with AI models before deployment. This ‘wargame’ approach allows companies to test AI decision-making in a risk-free environment, ensuring that what works in theory translates into real-world performance. The goal is to avoid costly failures in critical moments, especially in industries where trust and integrity are everything.

AI success isn’t just about spotting crises or generating convincing chats. It’s about finishing what it starts, reading critical internal data, and maintaining discipline under pressure. Firmulate’s live experiment proves that even the most thorough AI models can stumble if they neglect prioritization and ethical discipline—lessons that matter for any industry relying on AI-driven decisions.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.