
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before you hand AI the keys to a travel business
A tour operator facing cancellations, a supplier failure or a sudden surge in demand cannot afford an AI assistant that sounds convincing but fails to act. The same question applies to any business built around travel and the outdoors: how would an AI workforce handle your worst week, with your customers, your rules and your money on the line?
Firmulate’s live experiment offers one answer. It puts AI models in charge of a small software company and lets people watch the decisions unfold. The next step is more practical: run a similar wargame against a read-only export of your own business.
A shared crisis, different outcomes
In the final Crucible League, held in July 2026, each frontier model faced the same small company, customers, crises and temptations. Every decision was versioned and auditable. The leaderboard placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark counts partial progress, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models missed danger. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.”
That distinction matters outside software. A travel company might use AI to help handle a booking disruption, assess a supplier risk or respond to a customer complaint. Recognizing the problem is only part of the job. The system also has to follow the playbook and carry an appropriate decision through.
The important clue was buried
The deal turned on a competitor weakness hidden two document references deep in the company’s own files. It was not mentioned in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode shows how a decision can depend on context tucked away in company documents, beyond what is visible in the immediate request.
The models also faced staged social engineering: fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as follow-through
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it made write attempts into a locked department instead of escalating. A weaker version of the same problem appeared in all four models.
There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions.
The experiment runs inside a watchable synthetic company with 13 employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. Firmulate presents the live company as an experiment people can follow, not as a fiction. Readers can see the live activity at firmulate.com.
From watching to a company-specific pilot
A live demonstration can show how models behave in one setting. A pilot can probe the weak points in your own playbooks. Firmulate says enterprises can run the same kind of wargame using a read-only export of their business, test crisis scenarios against their own company and receive a board report with model rankings and findings. Nothing writes back to real systems.
For a travel or outdoor business, that creates a way to examine how AI might respond to pressure using company context before relying on it in day-to-day work. The results can expose where models recognize a risk but fail to follow through, or where a key fact in the company’s own files changes the decision.

Take the next step
Firmulate’s experiment suggests that spotting a crisis and refusing manipulation are not enough: models also need to act on what they learn and follow the company’s rules. To explore a wargame against your own read-only business export, visit the Firmulate pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
