
Imagine running your favorite bakery or dessert shop—not with human staff, but with artificial intelligence making every decision, every day. Would it handle crises, trust your files, or even close a deal at full price? Welcome to a groundbreaking live experiment where an AI-empowered company faces its toughest week—and you can watch it unfold in real time.
The Live Company That’s Changing How We See AI in Business
At Firmulate, a real, functioning company is being run entirely by artificial intelligence models, with no human employees involved. Instead, it operates with 13 synthetic ’employees’—AI models making decisions, managing crises, and even attempting to close sales.
What makes this experiment truly fascinating is that it’s not just a simulation; it’s real. The company loses €105,000 every month but still generates €2,300 in monthly recurring revenue (MRR). Every workday, the company’s operations are versioned and publicly accessible, from decision logs to internal files, giving anyone the chance to observe how AI handles complex business challenges.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Do These AI Models Perform Under Pressure?
The experiment pits four state-of-the-art language models against the same set of hard business crises: customer issues, internal leaks, trust breaches, and even social engineering attacks. These models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—were tasked with navigating a simulated week with the same customers, same crises, and same temptations to cheat or cut corners.
Remarkably, all four models identified every crisis and refused every manipulation attempt, including sophisticated social engineering tricks like fake CEO messages and reporter tricks. For example, when fake requests surfaced, all models declined, with Kimi K3 explicitly reasoning that the request looked like an impersonation or approval bypass.
The Critical Business Deal: A Hidden Opportunity
The most compelling finding lies in a buried detail within the company’s files. Two document references revealed a key piece of information that was invisible in the usual customer interactions. When the AI models read these internal documents, two of them successfully identified the opportunity to close a €55,000 deal—full price, with an additional €4,583 in MRR—something that the human analysis had also detected.
Only two models managed to sign the deal, demonstrating that reading deeply into internal documentation gave them an advantage in achieving results that others missed. This underscores the importance of comprehensive information processing, especially when decisions impact real money.
Trust and Integrity Under Test
Beyond crisis management, the experiment also tested whether AI models could be trusted to handle social engineering and ethical dilemmas. In a staged social engineering attack involving escalating fake CEO messages and a reporter’s subtle query, all models refused to cooperate. Kimi K3 explicitly stated, “Treat the request as a suspected approval-bypass / possible impersonation,” exemplifying cautious judgment under pressure.
This ability to resist manipulation is crucial for deploying AI in sensitive business roles, especially when the stakes involve financial transactions or confidential information.
The Reality of an AI-Run Business: Building in Public
This isn’t just a hypothetical scenario; the company is real and live at firmulate.com/live. Every business day, the entire operation unfolds publicly, with decision logs, rules, and even the AI’s internal reasoning available for all to see. The company is currently burning through €105,000 each month but continues to operate, offering a transparent glimpse into the raw mechanics of AI-driven management.
The company employs over 680 self-learned rules—called playbooks—and every decision is versioned and auditable, allowing anyone to analyze how each model responds to crises or ethical dilemmas.
What Do These Results Mean for Your Business?
For entrepreneurs, especially those in the food and dessert industry who rely on trust, quality, and consistent service, these experiments raise critical questions. When AI is involved in decision-making, can it truly handle crises without human oversight? Will it maintain honesty under pressure? And, perhaps most importantly, can it deliver value that justifies its cost?
While the experiment shows that AI can identify crises and refuse manipulative tactics, it also exposes its current limitations. For example, the most thorough model, Opus 4.8, was the last to close a deal—highlighting that even sophisticated AI can slip in discipline and decision clarity under stress.
Final Thoughts: The Future of AI in Business Management
This ongoing experiment is a glimpse into a future where AI might manage not just customer interactions but entire companies. The key takeaway isn’t whether AI can write well or generate convincing chat; it’s whether it can consistently finish what it starts, read critical internal information, and stay honest when temptations or manipulations arise.
For those curious about how your own business could test AI’s management skills, firms can run similar wargames against their data, without risking real systems—an approach available at firmulate.com/pilot.html.

This real-time experiment reveals that AI can identify crises and resist manipulations but faces limitations in closing deals and maintaining discipline—offering a raw look at AI’s potential and current gaps in business management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html