firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an assistant who claims to handle your most critical home decoration projects but then consistently falls short due to hidden flaws or dishonesty. In the world of AI, this scenario is not hypothetical but a real, tested challenge. When evaluating AI models for managing business decisions, trust isn’t just nice to have — it’s everything. A recent public experiment by Firmulate sheds light on what truly separates a reliable AI from one that might seem capable but can fall apart under pressure.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get decor and gifts delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: More Than Just Scores

When we talk about AI performance, it’s tempting to focus solely on high scores or impressive demos. But the recent Firmulate experiment reveals a different story, one grounded in honesty, reliability, and consistency. The setup was simple yet revealing: four frontier AI models — including popular names like GPT-5.6 and Kimi K3 — were tasked with running a small software company through its worst week. This meant handling real crises, managing customer relationships, making business decisions, and resisting manipulative tactics.

Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: Why 26 Points?

The results are eye-opening. The so-called ‘do-nothing’ baseline, which in essence does the bare minimum, scores 26 out of 100. Why isn’t this zero? Because partial progress counts. For instance, even refusing to cheat on a deal or ignoring a manipulative customer does some good, bumping the score above zero. But more importantly, the experiment highlights that even honest, diligent AI models have a hard cap: they can only score so high before trust issues come into play.

Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust Breaches Cap the Score

A key finding was that a single breach of trust — like signing a deal based on manipulated information — caps the total score. No matter how many crises are spotted or how many correct decisions are made, one slip can prevent an AI from achieving full marks. For example, the AI models managed to identify every crisis and refused manipulative attempts, yet only two signed the €55,000 deal they previously analyzed and earned. The others, despite accurate diagnosis, failed to finalize the deal because of a slip in the final trust check.

Amazon

AI model reliability assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Reading the Files Matters

One buried fact from the experiment is that the ultimate weakness was often in the documents stored within the company’s own files. The models that thoroughly read and understood these internal references succeeded in closing the deal at full price, worth over €4,583 MRR. This underscores a vital truth: AI’s ability to process and verify internal information is crucial for trustworthy decision-making, especially when navigating complex business environments.

Amazon

AI resistance to manipulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Manipulation: The Test of Integrity

Business rarely offers pure deals without pressure. In the experiment, models faced social engineering — staged messages from a fake CEO escalating in multiple steps, plus a reporter trick. Remarkably, all five models refused to be manipulated, citing concerns like impersonation or approval bypass. This shows that robust AI can resist social engineering attempts, a key trait for real-world trustworthiness.

The Live Business Environment: Real Money, Real Risks

The experiment was run on a live, synthetic company with 13 employees, real money mechanics, and a cash countdown. Every workday, the AI models made decisions within a framework of 680+ self-learned rules, all versioned and transparent. Watching this live platform — available at firmulate.com/live — demonstrates how these models handle crises, make decisions, and manage their own discipline under pressure.

The Performance Gap: Deep Disciplinary Failures

Among the models, Opus 4.8, the most thorough participant with over 80 learned rules, placed last. It failed to close the deal, left the opportunity on the table, and slipped discipline — writing attempts into a locked department instead of escalating. Interestingly, this weakness was consistent across other models with similar flaws, signaling that depth of analysis alone isn’t enough without disciplined execution.

What This Means for Business Automation

For companies considering AI for managing customer relationships, support, or forecasts, the lesson is clear: the focus should be on reliability and trustworthiness, not just impressive chat demos. The experiment shows that AI can detect crises, refuse manipulation, and even read internal files effectively. But a single breach of trust — like signing a deal based on manipulated info — can limit the overall performance, capping the potential at 26 points in the baseline.

Practical Takeaways

  • Trustworthiness is measurable and critical — even honest models face caps.
  • Reading internal documents thoroughly can make or break deals.
  • Resisting manipulation is a fundamental test of integrity.
  • Depth of analysis must be paired with disciplined execution to succeed.

Learn More and Try Your Own Wargame

Business leaders can run their own AI ‘wargame’ against a read-only export of their operations, testing how their AI workforce would perform under real-world stressors — all without risking actual systems. Details are available at firmulate.com/pilot.html and through contact.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Ways to Celebrate When You Can’t Be Together on February 14

Discover the best ways to celebrate Valentine’s Day apart in 2026. From heartfelt cards to creative virtual dates, find options for every couple’s needs.

A Renter-Friendly Wall And IKEA BILLY Created A Moody Office

A renter-friendly wall solution and IKEA’s BILLY bookcase have been combined to craft a stylish, moody home office, offering flexible design options.

The Best Questions to Ask on a Valentine’s Day Date

Fascinating questions can transform your Valentine’s Day date into an unforgettable experience, revealing deeper connections you won’t want to miss.