
Imagine trusting your future to an AI manager—one that not only makes decisions but also shows a distinct personality. Would it be reliable? Would it act ethically? As AI increasingly touches our lives, understanding how these models behave under pressure becomes crucial. Today, we take you inside a real, live experiment where several frontier AI models ran a functioning software company through its toughest week—every decision, crisis, and temptation laid bare. The question: which AI truly acts like a responsible manager?
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Putting AI Models to the Test
Firmulate, an innovative company specializing in AI management simulation, conducted a groundbreaking test with four leading AI models, including GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5. Each was tasked with running a small but real software firm, facing the same weekly crises: demanding customers, potential manipulations, and internal pressures. Every decision was recorded and made accessible for analysis—no faking, no editing. The goal: see which AI can manage honestly and effectively under stress.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Measure of Success
The results provided fascinating insights:
- All four models successfully identified every crisis and refused manipulation attempts, demonstrating solid ethical boundaries.
- Only two models managed to close a lucrative deal worth €55,000, earned through their own analysis and efforts. The other two hesitated, left the deal on the table despite recognizing the opportunity.
- A buried weakness was uncovered—models that read deeper into the company’s internal files were more successful at closing deals at full price, adding €4,583 MRR to the company’s revenue.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Personalities of AI Managers
Interestingly, the models showed distinct ‘personalities’ in how they handled social engineering and deception attempts. When fake CEO messages escalated over three stages, all models refused to participate—Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a shared cautiousness against manipulative tactics, reflecting a form of ethical guardrail.
As an affiliate, we earn on qualifying purchases.
The Real-World Business Mechanics
This experiment isn’t just about AI trickery. It’s set in a real operational environment—13 synthetic employees, real money mechanics, and daily decision-making that impacts a company’s bottom line. The company burns €105,000 monthly against a mere €2,300 in revenue, operating in a tense cash countdown. Every decision by the AI models influences the company’s survival, making this a true test of management qualities.
As an affiliate, we earn on qualifying purchases.
Comparing the Models: Who Did Best?
Here’s how the models scored:
- gpt-5.6-sol scored 95 out of 100, successfully closing a full-price deal and uncovering critical information buried in internal files.
- Kimi K3, a newcomer, scored 93, also closing the deal with the cleanest discipline of the field.
- Sonnet 5 scored 88, closing the deal but with a few process slips.
- Fable 5 scored 77, also closing the deal but with more slips in discipline and process.
It’s notable that the highest-scoring models were those that read deeper and stayed disciplined under pressure, revealing their management ‘personality’ traits.
What Does This Mean for Your Business?
As AI becomes more integrated into daily operations—handling customer relations, support queues, or financial forecasting—the key question isn’t if they can write well. Instead, it’s whether they can finish what they start, stay honest under pressure, and make decisions aligned with your values. The models’ personalities influence their management style: some are thorough and cautious, others terse and efficient.
Try It Yourself
If you’re curious how your current AI tools would perform, you can run the same wargame against your business data. This is a safe, read-only simulation that helps you evaluate management qualities without risking real-world consequences. Learn more at firmulate.com/pilot.html. And for a fun challenge, test your intuition with the ‘Guess the Model’ quiz at firmulate.com/quiz.html.

In the evolving landscape of AI management, personality, discipline, and integrity matter as much as technical prowess. Real-world tests reveal which models truly act like responsible leaders—and which need further development. Before trusting AI with your business, see how it handles crises, manipulations, and opportunities in a safe, real-time simulation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.