Every few weeks a new model arrives that is supposedly better at exactly the thing you use a model for. The cheap way to find out is to switch, play with it for an evening, and decide it "feels sharper." That is how you end up trading one set of mistakes for another without noticing which ones you bought.
This note is about a comparison I ran in a private game project, where a model reads a player's free-text instruction and turns it into an ordered plan that the game engine then checks, prices and executes. The model never changes the world directly. It proposes; the engine decides. That boundary is what makes a fair comparison possible at all, because the engine can tell you exactly what a proposal would have done.
Freeze everything except the model
The test set is twenty situations — physical, social, investigative and mixed — each with two alternative requests and a paraphrase of the first. Sixty requests in total. For each one, the expected plan was written down before any model saw it, including the exact cost the engine should quote.
Then everything is frozen: the source files and their hashes, the full request bytes sent to the provider, the game state each request runs against, the random seeds. Each request gets exactly one slot. There are no rerolls. If a run is interrupted, resuming reads back the stored answer instead of asking again, and a replay of the whole batch makes zero new calls.
That last property sounds like bookkeeping, but it is the point. If you can re-ask, you will — and the score quietly turns into "best of three."
What the numbers said
The incumbent model matched the expected plan on 51 of 60 requests. The candidate matched 57 of 60. Every request the incumbent got right, the candidate also got right; the six gains were all new passes, not trades. On the paraphrase pairs — the same request worded differently — the candidate produced the same plan for 18 of 20 pairs against 16.
The more useful finding was in how it failed. The incumbent's worst habit was adding work nobody asked for: asked to gather evidence and reach an agreement, it also quoted a repair job that cost an extra hour. The candidate stopped doing that. Its remaining misses were mostly cases where the request was genuinely ambiguous.
It was not free. Median latency went from 1.3 seconds to 5.6, and the 95th percentile from 1.4 seconds to 18 — uncomfortably close to the 20-second timeout per attempt. Token usage rose by about 18%. And some problems survived the swap unchanged: an internal reference alias leaking into player-facing text, and physical searches described as if they had produced social facts.
Prompt changes moved the score more than I expected
Before the model comparison there were four rounds of changing the instructions with the same model. The scores went 53, 47, 45, then 51. The two dips came from changes I needed anyway — a privacy fix that stopped hidden context reaching the model, and a more explicit output contract — and both read like improvements. Without the frozen set I would have assumed they were neutral and never gone looking for what they cost.
This is also why the comparison had to use the latest instructions for both models. Comparing the new model on new instructions against the old model on old ones would attribute the instruction change to the model.
What changed, and what did not
- The candidate became the first choice for new requests, with the old model kept behind it.
- Fallback to the old model is limited to transport failures — timeouts, outages, quota. A bad answer is never a reason to quietly ask the other model for a better one. That would hide exactly the errors the test exists to count.
- Every stored answer records which model actually produced it, so a fallback answer is never credited to the new model.
- The engine-owned parts — the ordered steps, the full cost, the confirmation before anything risky — stay prominent. 95% on a set I have looked at is not a reason to hide them.
The reviewer is a different model
Both the setup and the results were reviewed by a model from a different vendor before I trusted them: once before any request was sent, to check that only the model choice differed between the two batches, and once afterwards, to audit the artifacts and pair every case. It found nothing blocking before dispatch. After, it confirmed that all sixty payloads were byte-identical between the two runs, and it wrote down the defects the headline number hides.
The habit that carries over to any LLM feature: decide what "right" means before the model answers, keep one answer per question, and let something other than the model under test grade it.