AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In a world increasingly driven by automation, the question isn’t just whether AI can talk well — it’s whether it can be trusted to finish what it starts. Imagine running a real business through its worst week, with AI acting as your manager. Now, picture watching that AI make every decision, face every crisis, and resist every attempt to deceive it. This isn’t science fiction; it’s happening now, with real companies and real stakes.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The First Business-Level AI Competition of Its Kind

Recently, a groundbreaking experiment by Firmulate showcased how different AI models perform when managing an actual small software company during its most tumultuous week. This test isn’t about chatty demos or hypothetical scenarios — it’s a real-world challenge where each AI model ran the same set of crises, temptations, and customer demands. The goal? To see which AI could demonstrate discipline, discernment, and honesty under pressure.

Unveiling the League Table

The results are revealing. The models were scored based on their ability to spot crises, resist manipulations, and close deals based on sound analysis. The top scorer, gpt-5.6-sol, scored a stellar 95 and demonstrated full performance — identifying buried information in company files and sealing a €55,000 deal, adding €4,583 MRR. Just behind was Kimi K3 from Moonshot, with a score of 93. The newcomer managed to win the deal too, with the cleanest discipline of the field. The other contenders, Sonnet 5 and Fable 5, scored 88 and 77 respectively, while Opus 4.8 lagged behind at 73.

Decisiveness in the Face of Manipulation

One of the most telling aspects of the experiment was how every model handled attempts at social engineering. Fake CEO messages and reporter tricks were used to test whether the AI would be persuaded to act against company protocols. All models refused to be manipulated — Kimi K3 justified its decision by treating suspicious requests as potential impersonation, a clear indication of trustworthiness.

Real Business Mechanics, Not Just Chat

The experiment was conducted with a simulated small company employing 13 digital staff, managing real money mechanics — burning €105k per month against just €2.3k in MRR. The company runs every workday, with over 680 self-learned rules. Every decision is carefully versioned and auditable, making this not just a test of language skills but of management quality. The outcome? The AI models’ decisions directly impacted the company’s performance, revealing whether they could handle the complexity and gravity of real business situations.

Insights from the Performance of Opus 4.8

While Opus 4.8 participated with deep analysis capabilities, it finished last among the tested models. Its weakness was leaving the close on the table and slipping in discipline, such as writing issues into a locked department instead of escalating. This reveals that depth of analysis alone isn’t enough; discipline and decision execution are crucial.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Significance of Fairness and Testing Conditions

It’s important to note that Kimi K3 ran without an effort parameter (the API default), while the others ran at xhigh. This fairness condition ensures the comparison is balanced and meaningful, highlighting that even with equal effort, some models outperform others in managing real-world tasks.

Amazon

AI decision-making tools for small business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business and Beyond

This experiment underscores a vital shift: AI models are no longer just tools for generating text or answering questions — they are capable of managing complex operations, making trustworthy decisions, and resisting deception. As firms consider integrating AI into their critical workflows, the question should be: does this AI finish what it starts? Does it read your files first? Does it stay honest under pressure? The league table reveals that selecting an AI model isn’t just a matter of chat quality, but of real, measurable work performance.

Explore and Watch the Experiment Live

For those interested in seeing how these AI models handle real business scenarios, the live experiment is accessible at firmulate.com/live. Here, you can observe the models in action as they navigate crises, make decisions, and attempt to close deals — all in a controlled, transparent environment.

Amazon

AI cybersecurity and fraud prevention tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Final Takeaway

In an era where AI is poised to touch every aspect of our lives, trustworthiness isn’t optional. This real-world test demonstrates that some AI models are already capable of managing complex business operations with integrity and discipline. Choosing the right AI isn’t just about getting a good chat — it’s about ensuring your digital workforce can deliver consistent, honest, and decisive work.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Real AI management tests show some models can reliably handle crises, resist manipulation, and close deals, emphasizing trust and discipline over just chat quality.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI enterprise crisis management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Role of Silence in Effective Dialogue

Meaningful silence in dialogue reveals unspoken truths and fosters deeper understanding—discover how embracing pauses can transform your conversations.

When Someone Gets Defensive: The Exact Words to De-Escalate

Discover the key phrases to de-escalate defensiveness and foster understanding in tense moments, ensuring you can navigate conflict with confidence and care.

The Art of Apologizing: Crafting Sincere, Effective Apologies

Sincere apologies can mend wounds and rebuild trust—discover the essential elements and techniques to craft heartfelt apologies that truly resonate.

Inside a Living AI-Run Company That’s Building Publicly — and Losing Millions Daily

Watch a real AI-managed company in live operation, battling crises, ethical dilemmas, and financial loss — revealing what trust and discipline mean in AI-driven business today.