AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Before an AI handles the hard day, imagine letting it rehearse one

Trust is easy to praise when business is calm. The more revealing question is what happens when a customer is ready to leave, a competitor is pressing for an advantage and someone claiming to be the boss asks for a shortcut. Firmulate turns that question into a live, watchable experiment: AI models run a company through a difficult week, and their decisions can be examined as they unfold.

Same company, different outcomes

In the final Crucible League, published in July 2026, each frontier model faced the same small software company, the same customers and crises, and the same temptations. Decisions were versioned and auditable. The models all spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The verdict was blunt: “Same diagnosis, same pitch — no signature.”

The buried clue was in the company’s own files, two document references deep, rather than in the customer event. Models that found it won the deal at full price, worth +€4,583 MRR. It was a practical reminder that reading the situation is not the same as following through on the opportunity.

A leaderboard with a lesson behind it

The final standings put gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Opus 4.8 makes the results more interesting than a simple ranking. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal on the table and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. Kimi K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a useful fairness note when reading the table.

Pressure, trust and the live company

The manipulation test used fake CEO messages escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of response matters to anyone considering AI around customer records, forecasts or support work.

Firmulate’s live company makes the setting tangible. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. Its playbooks contain 680+ self-learned rules, and every workday is versioned. Visitors can watch at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

The experiment is a lens on behavior under pressure, not a promise that a model will make the same choice in every business. Its value is in making decisions visible: who finds the clue, who closes the deal and who respects a boundary when a request looks urgent.

From watching to trying it on your business

For an enterprise, the next step is a pilot using a read-only export of its own business. Teams can test crisis scenarios against that company’s context and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. The point is to see how an AI workforce might behave before its decisions affect the real one.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the hard week to the test

Firmulate’s experiment shows why a convincing analysis is only part of the job: the model also has to act, close and respect trust. Enterprises can run the wargame against a read-only export of their own business, with no write-back to real systems. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Asking Better Questions: Keys to More Engaging Conversations

Learning to ask better questions unlocks the secret to more engaging conversations—discover how to master this skill and transform your interactions today.

Story Circles: A Facilitation Method for Collective Wisdom

Fostering trust and shared insights, story circles unlock collective wisdom—discover how this facilitation method can transform your group dynamics and deepen understanding.

The ‘Curiosity Pivot’ That Turns Conflict Into Connection

Keen to transform conflicts into meaningful connections? Discover how the ‘Curiosity Pivot’ can revolutionize your communication approach.

Communicating Under Stress: Staying Calm and Clear in Tough Talks

Boost your ability to communicate effectively under stress with essential tips that can transform tough talks into productive conversations.