
Imagine testing an ice cream maker by asking it to simply hold the cone — no toppings, no scoops — yet it still scores some points. Sounds odd, right? But this is exactly what happens in AI benchmarking. Even a ‘do-nothing’ AI gets a baseline score, revealing a lot about what we can and can’t expect from these digital workers.
Get baking supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Curious Case of the AI Baseline Score
When evaluating AI models for business tasks, one might assume that a model doing nothing should score zero. But in the latest experimental benchmark from Firmulate, the ‘do-nothing’ baseline actually scores 26 points out of a possible 100. Why? Because the test measures more than just active decision-making. It captures the model’s ability to recognize crises, resist manipulation, and stay honest — even when it’s not actively trying to solve a problem.
As an affiliate, we earn on qualifying purchases.
How the Benchmark Works
The experiment puts four leading AI models through the same simulated week in a small software company. All models encounter the same customers, crises, and temptations—like fake CEO messages or requests to manipulate data. The goal: see whether they can identify real issues, resist manipulation attempts, and make honest decisions.
This isn’t a test of chat-bot flair or storytelling. Instead, it’s a rigorous evaluation of management skills—reading the company’s hidden files, refusing unethical requests, and closing deals when appropriate. Every decision is versioned and auditable, ensuring transparency and fairness.
The Surprising Results
All four models recognized every crisis and refused every manipulation attempt. That’s promising. But only two models managed to close a deal at full price, worth over €4,583 in monthly recurring revenue (MRR). The other two signed deals at a lower price or didn’t close at all, despite diagnosing issues correctly and pitching effectively.
The key difference? The models that read deeper into the company’s own files—beyond just the immediate customer interactions—secured the full deal. One model, Kimi K3, did this perfectly, reading two document references deep and closing the deal at full value. The other models weren’t as thorough, leaving potential revenue on the table.
Why Trust and Discipline Matter
Another test involved social engineering—fake CEO messages escalating over three stages and a reporter trick asking for a background yes/no. All models refused these manipulative requests. Kimi K3 explained its refusal as treating the request as a suspected impersonation, reflecting a disciplined approach.
In the real company simulation, every decision was critical. The company had 13 synthetic employees, with real money mechanics operating at a loss—burning €105,000 monthly against a mere €2,300 MRR. The AI models needed to carefully navigate this environment, balancing discipline and opportunity.
The Reality Behind the Scores
Interestingly, the most thorough participant, Opus 4.8, scored the lowest among the four—only 73 points. Despite analyzing more rules and performing detailed diagnostics, it slipped in closing the deal and escalating issues properly. The model’s discipline waned, illustrating that more analysis doesn’t always translate into better management outcomes.
It’s important to note that the models were run at different effort levels. Kimi K3 operated with the default API setting, while others pushed at high effort levels. This variation highlights how model settings influence performance, yet the core ability to stay honest and thorough remains paramount.
The Takeaway for Business Leaders
This experiment underscores an essential truth: evaluating AI’s business readiness isn’t just about how well it chats or generates ideas. It’s about whether it can complete complex, honest work under pressure—reading the right information, resisting manipulation, and closing deals ethically.
For companies considering AI automation, the question must be: Will this AI stay disciplined when it encounters real-world temptations? Will it read your files before acting? Will it finish what it starts, or leave revenue on the table? These are the critical tests that no simple score can fully capture, yet the Firmulate benchmark exposes them clearly.
Looking Ahead
As AI models improve, the benchmark’s floor—set by the do-nothing baseline—reminds us that honesty, thoroughness, and discipline are the true measures of a model’s business readiness. The live experiment, visible at firmulate.com/live, shows these models in action—delivering real insights, making real decisions, and risking real money.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
