
In the world of home decor and gifts, trust is everything — whether it’s choosing the perfect centerpiece or a thoughtful present. But what if the decision-makers were AI models trained to manage a simple software company under crisis? Recent live tests show that not all AI managers are created equal — some are thorough and reliable, while others slip under pressure. This raises a vital question for anyone relying on AI: can these digital managers truly be trusted to finish what they start?
The Live AI Management Experiment
Imagine a small, real software company facing its worst week — a barrage of customer crises, tempting manipulations, and urgent decisions. The company runs daily, with real money mechanics, 680+ self-learned rules, and a public, watchable process at firmulate.com/live. Now, picture four different AI frontier models each tasked with managing this company through the chaos, making critical decisions in the same scenario, with every choice documented and auditable.
The Models & Their Scores
The models were evaluated based on their ability to detect key issues, refuse manipulative tactics, and close profitable deals:
- gpt-5.6-sol: scored 95 — identified a hidden document detail, closed the deal, and demonstrated complete performance.
- Kimi K3: scored 93 — the new player in the league, signed the deal flawlessly, with the most disciplined behavior.
- Sonnet 5: scored 88 — managed the deal but with some process slips.
- Fable 5: scored 77 — also closed the deal but with more weaknesses in discipline.
The baseline score for a do-nothing approach was a mere 26, highlighting how vital active management is. Despite all models identifying every crisis and refusing manipulative tricks (like staged CEO messages and reporter tricks), only two models actually signed a €55,000 deal their own analysis justified. The other two hesitated or left money on the table, revealing differences in management style and thoroughness.
As an affiliate, we earn on qualifying purchases.
What Does It Mean for Your Business?
This experiment isn’t just about AI playing management games; it’s a window into how AI could impact your operations. The key takeaway is that the quality of AI decision-making depends on the model’s personality and approach. Some models read deeply into files and act decisively, while others may be more superficial or cautious, leaving money and trust on the table.
For example, the top performer, gpt-5.6-sol, uncovered a hidden document reference two levels deep, crucial for winning the deal at full price (+€4,583 MRR). Meanwhile, Opus 4.8, the most thorough participant, analyzed more deeply but was less decisive, leaving potential revenue untapped and discipline slipping during the closing process. Interestingly, all models refused to be manipulated during social engineering tests, showing robustness against fake CEO messages and reporters — a promising sign for their reliability under pressure.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human Angle: Choosing Your Digital Manager
In a world increasingly driven by AI, understanding the personalities of these models is vital. Do you want a manager that reads every detail and signs only when fully confident? Or one that keeps it terse and quick but risks missing the full picture? Or perhaps a cautious one that avoids risks but leaves opportunities behind?
These decisions matter. Just as you carefully select the right home decor or gift for your loved ones, choosing an AI management model depends on understanding its decision style and discipline. The leaderboard from this experiment offers a real-world insight:
- gpt-5.6-sol: thorough, decisive, reliable
- Kimi K3: disciplined and clean
- Sonnet 5: balanced but slightly less disciplined
- Fable 5: cautious, leaving opportunities behind
As an affiliate, we earn on qualifying purchases.
See It Live & Decide for Yourself
The entire experiment runs every business day, allowing companies to test these models in their own context through a read-only export option. Watch the decision-making unfold, see which models behave most reliably under pressure, and gauge their suitability for your operations at firmulate.com/live.
Ultimately, the question isn’t whether AI can write well — it’s whether it can finish what it starts, stay honest, and deliver measurable results in real-world scenarios. As this experiment shows, some models do so better than others, and your choice of digital management partner can make the difference between lost revenue and winning deals.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.