AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Can you trust a good intention when the moment to act arrives?

Faith traditions often ask us to look beyond what someone says they believe and notice what they do when the choice becomes difficult. A live experiment from Firmulate puts a practical version of that question to artificial intelligence: when an AI can recognize a crisis and explain the right move, will it actually follow through?

Same crises, different choices

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Their decisions were versioned and auditable, making it possible to examine what happened rather than judge a polished answer in isolation.

The league’s final standings put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment also treated a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”

Seeing the right answer was not the same as doing it

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. They reached the same diagnosis and made the same pitch, but some never put a signature on the result: “Same diagnosis, same pitch — no signature.” That gap between judgment and follow-through is easy to miss in a chat demo.

The decisive clue was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The result suggests that reliable action depends on more than sound instincts: an agent also has to look carefully at the information already available to it.

Pressure, trust and uneven discipline

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a more complicated portrait. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The deal was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. Careful analysis, then, did not guarantee careful execution.

There is a fairness detail for readers interpreting the ranking: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate’s live company adds another dimension to the experiment. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and the live experiment is watchable at firmulate.com.

Readers can also explore 242 real, unedited management decisions in Firmulate’s “guess the model” quiz at firmulate.com. The exercise makes the broader point tangible: evaluating an AI workforce means looking at decisions under pressure, not just the confidence or fluency of its explanations.

From watching to trying it in your own business

For an enterprise, the next step is a pilot against a read-only export of its own business. The wargame can put company-specific crisis scenarios to work and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The question shifts from whether an AI can describe a company’s values to how it behaves when those values meet a hard choice.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Watch the live experiment, then bring the same kind of scrutiny to your own business. Explore a Firmulate enterprise pilot using a read-only export of your company; contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Heated Blankets Work So Well for Comfort-Focused Nights

Discover the best heated blankets for comforting self care in 2026. Find top picks for warmth, safety, and ease of use tailored to your needs.

This Bath Upgrade Can Make Self-Love Nights Feel Extra Special

Discover the top luxury bath trays for self love rituals in 2026. Find the perfect blend of style, function, and relaxation with our curated picks.

How to Recognize Healing Burnout Before It Takes Over

Learning to recognize healing burnout early can prevent its takeover, so understanding the signs is crucial to maintaining your well-being.

Mukhi Rudraksha Surges In Global Coverage

Mukhi Rudraksha has seen a sharp increase in international media mentions, with GDELT reporting 36 times the baseline coverage recently.