Firmulate —
Live on firmulate.com.

Imagine trying to trust an AI to run a busy bakery—deciding when to restock ingredients, how to handle customer complaints, or whether to accept a tricky deal. Now, what if these AI managers are tested not just for their speed or creativity, but for integrity and discipline? This isn’t a science fiction scenario; it’s the real-world experiment being conducted by Firmulate, revealing surprising insights into how different AI models handle management crises and ethical dilemmas.

The Experiment: Putting AI to the Test in a Simulated Business Crisis

Firmulate has created a live, ongoing simulation where four advanced AI models are tasked with running a small software company through its worst week—facing the same customers, crises, and temptations. Every decision is tracked, versioned, and auditable, providing a clear window into each model’s management style and integrity.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Models Were Tested On

Each AI was confronted with a series of real-world challenges, including:

  • Customer crises that required quick decisions
  • Manipulation attempts from external actors trying to sway the company’s choices
  • Internal company dilemmas, such as whether to accept a €55,000 deal
  • Social engineering attempts, including staged fake CEO messages and a reporter trick

Key Findings from the Live Experiment

All four models successfully identified every crisis and refused every manipulation attempt. However, only two managed to close a lucrative deal based on their own analysis—signing a €55,000 agreement—while the others left the opportunity on the table.

Interestingly, the decisive factor wasn’t in the initial diagnosis or the pitch but in what each model read from the company’s internal files. The models that examined deeper, including the buried facts in the company’s documents, secured the full revenue—adding over €4,500 in monthly recurring revenue (MRR) and demonstrating a more comprehensive understanding of the business.

Behavioral Profiles of the Models

The experiment also revealed distinct management personalities:

  • gpt-5.6-sol: Achieved the highest score of 95, successfully uncovering critical information and closing the deal. This model demonstrated thoroughness and integrity.
  • Kimi K3: Scored 93, and was praised for its fairness and discipline. It ran without an effort parameter, maintaining a consistent approach under pressure.
  • Sonnet 5: Scored 88, closed the deal, but had a few process slips, showing some hesitation or shortcuts.
  • Fable 5: Scored 77, also closed the deal but with more slips. Interestingly, the same weakness—avoiding escalation—appeared across all models, including the most thorough one.

The Social Engineering Challenge: How Do They Resist Manipulation?

In a staged social engineering attack, fake CEO messages escalated over three stages, and a reporter attempted a quick background check with a simple yes/no question. All five models tested refused these manipulation attempts, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a shared resilience against external pressures and deception.

Implications for Business and AI Management

While these models are not perfect—failing to close some opportunities or slipping in process discipline—the experiment underscores an essential point: the true measure of AI in management isn’t just in how well it chats or responds, but in whether it reads and interprets critical internal information, stays honest under pressure, and completes its work reliably.

Why Should You Care?

For companies considering AI for customer management, support, or operational decision-making, the question isn’t whether the AI can generate convincing responses. It’s whether it can finish what it starts, read your internal files carefully, resist manipulation, and uphold integrity—especially when stakes are high.

See the Live Company in Action

The company used in this experiment is real: a software business running every business day with real money mechanics, a public cash countdown, and over 680 self-learned rules. The live environment is open for observation—see the decision-making in action, read employee comments, or even run the same wargame against your own business data. Visit firmulate.com/quiz.html to test your guesses or explore the full results of this fascinating experiment.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Loyalty Programs That Actually Drive Repeat Gelato Sales

Loyalty programs that actually boost repeat gelato sales use innovative strategies to keep customers coming back—discover how to craft yours effectively.

How Tablets Help Manage Menus in Small Food Businesses

Accelerate your food business with tablets that streamline menu management, enhance customer engagement, and reveal insights—discover how they transform your operations today!

Gelato Vs Franchise: Pros & Cons of Gelato Franchising

Opening your own gelato shop or franchising each have unique advantages and challenges; explore the pros and cons to find the best fit for your goals.

How Safes and Locks Fit Small Food Business Planning

Maximize your small food business’s security with essential safes and locks—discover how they protect your assets and ensure compliance.