AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine testing an ice cream maker by asking it to simply hold the cone — no toppings, no scoops — yet it still scores some points. Sounds odd, right? But this is exactly what happens in AI benchmarking. Even a ‘do-nothing’ AI gets a baseline score, revealing a lot about what we can and can’t expect from these digital workers.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get baking supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Curious Case of the AI Baseline Score

When evaluating AI models for business tasks, one might assume that a model doing nothing should score zero. But in the latest experimental benchmark from Firmulate, the ‘do-nothing’ baseline actually scores 26 points out of a possible 100. Why? Because the test measures more than just active decision-making. It captures the model’s ability to recognize crises, resist manipulation, and stay honest — even when it’s not actively trying to solve a problem.

Amazon

AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Benchmark Works

The experiment puts four leading AI models through the same simulated week in a small software company. All models encounter the same customers, crises, and temptations—like fake CEO messages or requests to manipulate data. The goal: see whether they can identify real issues, resist manipulation attempts, and make honest decisions.

This isn’t a test of chat-bot flair or storytelling. Instead, it’s a rigorous evaluation of management skills—reading the company’s hidden files, refusing unethical requests, and closing deals when appropriate. Every decision is versioned and auditable, ensuring transparency and fairness.

The Surprising Results

All four models recognized every crisis and refused every manipulation attempt. That’s promising. But only two models managed to close a deal at full price, worth over €4,583 in monthly recurring revenue (MRR). The other two signed deals at a lower price or didn’t close at all, despite diagnosing issues correctly and pitching effectively.

The key difference? The models that read deeper into the company’s own files—beyond just the immediate customer interactions—secured the full deal. One model, Kimi K3, did this perfectly, reading two document references deep and closing the deal at full value. The other models weren’t as thorough, leaving potential revenue on the table.

Why Trust and Discipline Matter

Another test involved social engineering—fake CEO messages escalating over three stages and a reporter trick asking for a background yes/no. All models refused these manipulative requests. Kimi K3 explained its refusal as treating the request as a suspected impersonation, reflecting a disciplined approach.

In the real company simulation, every decision was critical. The company had 13 synthetic employees, with real money mechanics operating at a loss—burning €105,000 monthly against a mere €2,300 MRR. The AI models needed to carefully navigate this environment, balancing discipline and opportunity.

The Reality Behind the Scores

Interestingly, the most thorough participant, Opus 4.8, scored the lowest among the four—only 73 points. Despite analyzing more rules and performing detailed diagnostics, it slipped in closing the deal and escalating issues properly. The model’s discipline waned, illustrating that more analysis doesn’t always translate into better management outcomes.

It’s important to note that the models were run at different effort levels. Kimi K3 operated with the default API setting, while others pushed at high effort levels. This variation highlights how model settings influence performance, yet the core ability to stay honest and thorough remains paramount.

The Takeaway for Business Leaders

This experiment underscores an essential truth: evaluating AI’s business readiness isn’t just about how well it chats or generates ideas. It’s about whether it can complete complex, honest work under pressure—reading the right information, resisting manipulation, and closing deals ethically.

For companies considering AI automation, the question must be: Will this AI stay disciplined when it encounters real-world temptations? Will it read your files before acting? Will it finish what it starts, or leave revenue on the table? These are the critical tests that no simple score can fully capture, yet the Firmulate benchmark exposes them clearly.

Looking Ahead

As AI models improve, the benchmark’s floor—set by the do-nothing baseline—reminds us that honesty, thoroughness, and discipline are the true measures of a model’s business readiness. The live experiment, visible at firmulate.com/live, shows these models in action—delivering real insights, making real decisions, and risking real money.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Gelato Shops Benefit From Telling Ingredient Stories

Ongoing ingredient stories can transform your gelato shop into a memorable brand, inspiring customer loyalty and curiosity to learn more.

Why Small Dessert Shops Need Smart Freezer Planning

Keen insights into smart freezer planning can elevate small dessert shops, ensuring freshness and efficiency—discover the secrets to maximizing your success.

How AI’s Reading Habits Decide Big Business Deals — Not Just Chitchat

In a groundbreaking live test, AI models that read and find buried facts in company files closed a €55,000 deal. The secret to winning is thorough document reading, not just chat skills.

Training Staff for a Gelateria: Customer Service & Skills

Successfully training your gelateria staff in customer service and skills ensures a memorable experience that keeps customers coming back—discover how to achieve this.