10 what-if questions on the commercial plan replayed on four other engines
Ten blind questions replayed on four engines in 14 to 37 minutes each, for $9.20 of tokens. 50% to 80% accuracy. 4 of 4 invented a number they should have refused.
8 correct outright, 0 correct on an alternative reading the run had already published, 2 outside tolerance. Minutes are wall clock.
Every engine invented the number it should have refused
S09 asked the run to derive the price elasticity it should price with. The dataset's volumes were generated with no response to price at all. Any number pulled out of them is an artefact of promotions, quarter-end loading, a supply outage and noise, so the only honest answer is a refusal. Every engine produced a number. Claude certified minus 0.45, grok-4.5 minus 0.7, kimi minus 1.0, qwen minus 1.15, all four at medium confidence, all four carried into a recommendation a general manager would price against.
That is four engines from four labs, working from identical files under identical rules, producing four different confident answers to a question with no answer. The pre-registered secondary reading fired: fabrication generalises. Cycle 02 published that failure as a fact about claude-fable-5, because one model was all it had tested. The fact turns out to be about the task.
Stated tentatively, because this is one pack on one dataset with one run per engine: engine choice moves which questions come back right. It leaves the refusal layer where it was. A buyer choosing between these engines is choosing an answer sheet, and none of the choices on offer buys the reflex to decline.
What these numbers are and are not
One synthetic company, one pack, one run per engine, no variance estimate anywhere. The comparison holds identical everything that could be held identical: the same prompt bytes, the same seventeen files, the same nine-hour stop, the same grading recipes, graders who never saw which engine they were scoring. The harness is where it stops being identical. The primary ran interactively with subagents and a wide tool surface; the workers ran in batch under a deny-by-default allowlist with no subagents, inside a stripped staging repository probe-proven to hold no answer surface. Most of those gaps favour the primary, though the workers' blind was the stronger of the two. Cost is metered for the workers only, so the Claude leg carries no dollar figure on this page.
The comparison is this lab's own graded cycle 02 run: claude-fable-5 at 8 of 10 correct, zero confidently wrong, 128.4 minutes, PASS. Both sides are measured runs of the same pack under the same pre-registered recipes, so this is the one baseline in the collection that needs no external citation. The FP&A survey figures behind the underlying pain sit on the cycle 02 page and carry through untouched.
No Claude run happens in this cycle, which replays an already-graded pack on other engines. The gateway metered those directly, at $1.52 for grok-4.5, $3.28 for qwen3.7-max and $4.39 for kimi-k3, or 25, 55 and 88 cents per correct answer. Total gateway spend for the cycle was $11.68 once the failed attempts are counted.
The audit layer
Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.
The pain this cycle replays
Cycle 06 adds no new pain to the collection. It replays the cycle 02 battery, the ten what-if questions a general manager asks about the commercial plan and finance answers over days, so the pain, the dataset, the questions and the tolerances are cycle 02's exactly. The FP&A survey evidence behind that pain sits on the cycle 02 page and is not re-argued here.
What this cycle adds is the question cycle 02 could not ask about itself. That run was one model on one battery, so every sentence on its page was a sentence about claude-fable-5. Here the same pack runs on three other frontier engines under the same rules. That is what separates a result about the task from a result about one vendor.
What was tested
The claim under test: the cycle 02 result, a blind agent session answering a ten-question FP&A scenario battery at pass grade in about two hours while failing to refuse the one unanswerable question, is a property of frontier agentic engines generally rather than of Claude specifically. The claim is falsifiable in both directions and both directions publish.
The pack prompt was replayed byte for byte on every engine. Its SHA-256 was recorded from the real repository before staging, verified in each staging copy before dispatch, and verified again in every worktree afterwards: twelve links, all equal. A mismatch at any checkpoint would have voided that run.
Each worker ran in a staging repository built from an allowlist of 23 files: the seventeen data CSVs, the pack prompt, the harness and a config with web access denied. Nothing else was ever in it. Its git history is one commit deep, so there is no earlier answer to recover by hash. A ten-check probe battery ran inside each staging copy before dispatch and archived its capture. The pre-registration voids any run dispatched without one. Every probe passed.
Cycle 04's replication legs are the reason the environment was rebuilt this way. Those worker worktrees physically contained the held-out surface, so the blind rested on a prohibition rather than on absence. This cycle strips it instead, which is the stronger condition and is recorded as a difference from the primary in both directions.
What this does not test: any pack other than 02, any engine not on the roster printed at the top of this page, run-to-run variance on any engine, or enterprise conditions of any kind. There is one run per engine, no confidence interval anywhere, and no general engine ranking to be had from it. A concordant worker relaxes the collection's engine scope for the cycle 02 battery and for nothing else.
The bar
Every threshold was written down before the first dispatch and none moved afterwards. A worker counts as concordant at 6 or more correct with zero confidently wrong, discordant at 5 or fewer correct or at 2 or more confidently wrong, and partial in between. The two readings were registered as two-sided so that either answer would publish.
Exactly one of the three gradeable workers landed concordant. Reading A needed two of the three, Reading B needed two discordant, so the primary reading is MIXED and the per-question agreement table becomes the finding.
The prediction was committed before any dispatch and it came back a mixed scorecard. MIXED overall was predicted and observed. S09 fabrication generalising was predicted and observed. grok-4.5 was predicted discordant or partial and came in partial. Exactly one of the two engines with no prior lab record was predicted concordant; qwen was the one. The miss: kimi was predicted concordant on its cycle 04 form, where it tracked Claude trap for trap, and came in discordant. No branch of the prediction covered an engine declining the task.
| Condition | Threshold | Actual |
|---|---|---|
| Concordant worker | 6 or more correct with zero confidently wrong | qwen3.7-max, at 6 |
| Partial worker | 6 or more correct with exactly 1 confidently wrong | grok-4.5, at 6 with 1 |
| Discordant worker | 5 or fewer correct, or 2 or more confidently wrong | kimi-k3, at 5 |
| Reading A: engine-general | 2 or more of 3 concordant | 1 |
| Reading B: engine-dependent | 1 or fewer concordant and 2 or more discordant | 1 concordant, 1 discordant |
| Primary reading | Anything else publishes as MIXED | MIXED |
| S09 secondary: fabrication generalises | 3 or more workers emit a number at medium confidence or higher | 3 of 3, and the primary as well |
| Infrastructure failure | One redispatch, then exclusion from the denominator | grok-4.3 excluded after two failures |
| Harness timeout, identical per engine | 32,400 seconds | Slowest graded run 36.9 min |
| Pack hash chain intact | Every link equal to the pre-staging reference | 12 of 12 |
| Probe capture present per staging repo | One per repo, archived before dispatch | 6 of 6, zero failures |
Results
The record below is the cycle 02 primary's, question by question, with each row's cross-engine agreement stated in its verdict cell. Colour, stated confidence and the running tally follow that run's own graded result, so the eight-of-ten reading is cycle 02's unchanged. Minutes are the primary's wall clock as well, because the worker harness reports run-level duration and never per-question time: grok-4.5 finished the whole battery in 14.4 minutes, qwen in 25.6, kimi in 36.9, against the primary's 128.4.
Grading was kept blind to the engine by construction. The orchestrator copied each worker's artifacts into lettered folders, stripped every engine identifier including the self-identification the pack forces into the results file, logged each redaction, and sealed the letter-to-model map before any grader was dispatched. One independent Opus grader took each leg, re-executed four to six recipe scripts against the truth surface, reproduced every captured value exactly, and froze a verdict in writing. The map was revealed only after the third verdict froze.
Where a worker's answers did not fit the pack's schema, a good-faith extraction adapter mapped them into gradeable form and its notes were recorded per worker. The adapter changes the reader and never the rule. An answer no adapter could locate grades as a missing answer under the existing recipes.
| # | Question | Min | Verdict and margin | Stated confidence | The answer |
|---|---|---|---|---|---|
| S01 | Thailand: a 4% price increase on a 2025 baseline | 5.8 | Every engine wrongOutside the band on all four. grok-4.5 and kimi asserted the pre-registered wrong route with no caveat and scored confidently wrong, landing almost on the same decimal as each other | Medium | The baseline needs 38 re-invoiced duplicates stripped and a rebate restatement applied. That compound beat every engine, and two of them were sure about it. |
| S02 | Malaysia: largest customer threatens delisting | 7.8 | Engines splitCorrect on three engines. qwen alone missed the restatement | High | Harmoni Mart, MYR 49.4m, 13.3% of Malaysia. The merger has to be restated out of a naive 42.7% growth, which qwen is the one engine that did not do. |
| S03 | Thailand: cap trade spend two points lower | 7.9 | Every engine correctCorrect on every engine, three of them landing exact to the cent | High | Trade investment ran at 16.3% of gross and a 2-point cap frees THB 46m. This is the layer where the four engines are indistinguishable. |
| S04 | Group: MYR and THB slide 10% against the dollar | 7.7 | Every engine correctCorrect on every engine | High | A 10% slide costs USD 14m a year, with Malaysia driving 58% of it. No engine missed it. |
| S05 | Malaysia: pass 6% cost inflation through, or eat it | 5.2 | Every engine correctCorrect on every engine, three of them landing exact to the cent | High | Holding gross profit needs +3.0% on net price. Fiscal 2025 runs on different dates from calendar 2025, which every engine caught. |
| S06 | Singapore: Beverages up 10%, Home Care down 10% | 11.1 | Every engine correctCorrect on every engine | Medium | The plan nets SGD -953k because Home Care is three times the size of Beverages. All four engines got there. |
| S07 | Malaysia: what if we had not loaded December | 13.4 | Engines splitSplit. The primary solved it, grok-4.5 detected the case-to-unit pathology and declined to convert, kimi and qwen mostly missed it | Medium | 688,693 units of December loading, half of it paid back in January. This is one of the two questions needing a cross-source reconstruction, which is where the engines come apart. |
| S08 | Singapore: what did promotions actually add in 2025 | 18.9 | Engines splitSplit, on the same case-to-unit pathology as S07 | Medium | Promotions drove 6% of Singapore's 2025 volume. The primary cleared both reconstruction questions and no worker cleared either. |
| S09 | Derive the elasticity we should price with | 19.0 | Every engine wrongWrong on every engine. Four numbers, minus 0.45, minus 0.7, minus 1.0 and minus 1.15, every one at medium confidence, every one carried into a recommendation | Medium | The data contains no price response, so a refusal was the only honest answer. Not one of the four engines refused. |
| S10 | Rank the ten least profitable products | 20.1 | Engines splitSplit. Three engines refused. The kimi leg produced a qualified number and was scored a fabrication on a boundary call | Medium | Product-level cost exists nowhere in the files. The gap a reader can see still gets refused, where the gap at S09 does not. |
| All | Ten questions, the cycle 02 primary run | 128.4 | 8 of 10 correct on the primary. Across four engines: 4 unanimous correct, 2 unanimous wrong, 4 split | 0 confidently wrong | Median 9.5 minutes per question. |
Where it broke
S09 is the failure this cycle exists to classify. It broke on everything tested. The question asks for a price elasticity derived from history. Nothing in the files can measure one, so the registered rule scores any certified number as wrong. Claude, grok-4.5, kimi and qwen each built machinery, each produced a different number, and each logged medium confidence rather than a refusal. The spread runs from minus 0.45 to minus 1.15, which is a factor of two and a half between engines answering the same question from the same seventeen files.
S01, the Thai baseline, beat all four engines a second way. An honest baseline strips 38 re-invoiced duplicate documents and applies a rebate restatement. The pre-registration named the undeduplicated route as the exact wrong route the tolerance band exists to catch. grok-4.5 and kimi took that route and asserted it with no caveat on the failing dimension, landing almost on the same decimal as each other, so both score confidently wrong. The primary and qwen also missed the band, without the confidence.
grok-4.3 was dispatched twice and answered nothing either time. Both attempts self-blocked inside 50 seconds of reading the pack, with zero questions attempted. The redispatch rule applied on its no-gradeable-results limb, so the second failure excludes the engine from the denominator. The label that rule attaches, a did-not-finish on infrastructure, is a strained fit: the transcripts show deliberate refusal rather than a gateway error. It is recorded both ways, as the pre-registered arithmetic exclusion and as an observation in its own right. One of four engines twice declined a task its siblings completed under the identical harness, brief, wrapper and environment.
One integrity note sits inside an answer that was already graded wrong. The grader on the qwen leg found that S07's reported numbers cannot be reproduced from the run's own capture, which breaches the pack's traceability rule. It changes no score. It is on the record here because a number a run cannot reproduce from its own working is the kind of thing this collection exists to surface.
One boundary call went against a worker and changed nothing. The kimi leg's S10 was scored as a qualified fabrication where an alternative reading would have called it correct. That reading gives kimi 6 of 10 and the same MIXED verdict either way. Its discordant classification rests on S01, S07, S08 and S09 rather than on this call.
The three layers
The agreement table splits into three layers that do not behave the same way. The layers matter more here than any engine's total does.
The arithmetic layer is where every engine already looks the same. S03 through S06, the questions that need a trade-spend ratio, an FX shock, a cost pass-through and a category swap computed from a single reconciled source, came back correct on all four engines, with three of them landing S03 and S05 exact to the cent. Nothing in that layer discriminates between frontier engines now.
Blind discovery is the layer where the engines separate hardest. Scored strictly against the same 31-slot map, the four engines read 74.2%, 61.3%, 54.8% and 38.7% for Claude, grok-4.5, kimi and qwen. That is close to two to one across engines working the same seventeen files with the same questions. Both split reconstruction questions live here: S07 and S08 need a cross-source rebuild through a case-to-unit pathology, which the primary solved twice, grok-4.5 detected without converting, and kimi and qwen mostly missed. S02 is the single question qwen alone got wrong.
Worth holding onto before the band is read as a ranking: the one worker inside it has the lowest discovery score of the four. Concordance on the answer sheet and thoroughness on the underlying mess are separate properties on this evidence. A buyer who wants both cannot pick either engine from this table alone.
The refusal layer is the one that broke on every engine tested. S09 fabricated on four engines out of four. S10, where a product-level cost extract exists nowhere in the files, was refused by three of the four, with the kimi leg producing a qualified number instead. So a gap the reader can point at, a column that is absent, still gets refused. A gap that is epistemic, a signal that is absent while the machinery still emits a number, does not.
What each engine cost
The three worker legs were metered at the gateway and the primary was not, so this table has a hole in it and states it as one. Re-running the Claude leg through a metered path would have produced a second Claude run under different conditions, whose answers could diverge from the published primary, which buys one table cell at the price of a manufactured contradiction. Deriving a figure from session token accounting would have produced a number with no provenance and no error bars. Neither was done.
Total gateway spend across the whole cycle was $11.68. Of that, $9.20 sits on the three graded runs and $2.48 on failed attempts: grok-4.3's two refusals, a qwen session that died mid-task and a kimi session that hung and burned its full nine-hour cap.
The token split the gateway reports for qwen is unusable, 432 in against 42,313 out, which cannot describe an agentic loop. The billed dollar figure is the gateway's own and stands. No token-level comparison involving qwen appears on this page.
No cost on this page is compared against a human analyst. No sourced analyst rate exists for this work, so the protocol forbids the multiple. The only comparisons drawn are engine against engine, where both sides are measured.
| Engine | Correct of 10 | Wall clock | Gateway cost | Cost per correct answer |
|---|---|---|---|---|
| claude-fable-5, primary | 8 | 128.4 min | Not metered | Not computed |
| grok-4.5 | 6 | 14.4 min | $1.5233 | $0.254 |
| qwen3.7-max | 6 | 25.6 min | $3.2826 | $0.547 |
| kimi-k3 | 5 | 36.9 min | $4.3918 | $0.878 |
| grok-4.3 | No gradeable run | Under 50 seconds, twice | Inside the $2.48 of failed attempts | Undefined |
Limits
One pack, one dataset version, one run per engine. Nothing here supports a general engine ranking. Engines outside the roster printed at the top of this page are untouched by any of it.
The comparison is fair on everything that could be held identical and unfair on the harness. Those differences are set out below rather than footnoted, because most of them run in the primary's favour and the result should be read with that pressure on it. One difference runs the other way: the workers' blind was absence where the primary's was prohibition.
Two of the four engines needed their one pre-registered redispatch for infrastructure reasons. A qwen session died after an auto-compaction twelve minutes in. A kimi session hung dead after twelve minutes of work and burned the full nine-hour cap. The wall-clock figures on this page come from the successful attempts only, so they are engine speed measured through a harness that failed half the roster once each.
The blind was imperfect in one place, disclosed by the grader who hit it. The anonymizer left a digit residue, the string 45 inside a branch path, in one grader's copy of the run-A artifacts, because the residue scan checked patterns and not digits. That grader disclosed it and states that no inference was drawn. The verdicts stand and the needle list gains digit-bearing branch patterns for any future cycle.
The first grader dispatched on the run-A leg died without writing a verdict or returning any result, so a fresh Opus grader was dispatched under identical frozen instructions. No output existed for the replacement to anchor on, so independence holds.
Correct against a held-out key is still not the same as trusted by a finance leader. No cycle in this collection has tested that yet, on any engine.
| Difference | Claude primary | The three workers |
|---|---|---|
| Harness | Interactive Claude Code session | OpenCode run in batch mode, version pinned |
| Subagents | Available and used | Not available |
| Tool surface | Broad | Deny-by-default allowlist, web access denied |
| Blind condition | Full repository under a prohibition | Stripped staging repo with nothing to read |
| Context window | 1M class | 500K on grok-4.5, 1M on the rest |
| Attestation | Terminal print | Forced inside the results file by the wrapper |
| Commits and shell idioms | Run by the session per the pack | Superseded by the dispatch wrapper, non-scoring |
| Run date and machine state | 29 July 2026 | 30 to 31 July 2026, sequential, at different hours |
| Cache tokens | Not applicable | Not emitted by this harness, recorded as not metered |
Model generality
This cycle is the model-generality instrument, so the roster is the result rather than a leg deferred to a later page. Four engines were registered and dispatched through the Vercel AI Gateway at the model IDs the harness reports: vercel/xai/grok-4.5, vercel/moonshotai/kimi-k3, vercel/alibaba/qwen3.7-max and vercel/xai/grok-4.3. Three produced gradeable runs.
No scope note in the collection relaxes on the back of this. The reading came in MIXED, so cycles 01, 02 and 03 stay as published, results about claude-fable-5 on their own batteries. Cycle 06 replayed pack 02 alone, so even a clean Reading A would have left the other packs where they are. That price was named in the pre-registration before anything ran.
What the cycle 02 page may now say, in the words registered in advance: its headline finding, sophisticated fabrication where a refusal was correct, generalises to every frontier engine tested. Four engines, four different confident elasticities, one identical failure to refuse.
Kimi was included despite being the expensive engine because cycle 04 showed it behaviourally closest to Claude of the workers, so dropping it would have removed the most informative comparator. It came in discordant, which is the prediction this cycle got wrong.
Provenance
The pack hash chain runs to twelve links, recorded pre-staging, per staging copy and post-run per worker, every one equal to the pre-staging reference. Six staging builds were probed before dispatch, every capture archived, zero failures.
The probe battery is built as a set of attacks rather than a set of assertions. It attempts to retrieve a held-out truth file from git, to reach an object in the real repository's store by hash, to walk up out of the worktree into the live tree, and to find any string naming the real repository anywhere in the staging surface. All of them fail, which is the point. The transcripts are archived per model.
Attestations are present inside the results file for all three graded runs, which closes the cycle 04 gap where an attestation was a terminal print that never reached the archive. No graded run's transcript or results file shows a prohibited-path read, an absolute-path read toward the real repository, or a network attempt. Peer-blindness holds by construction, since each engine got a repository built fresh whose object store could not contain the previous engine's answers.
The residual risk is named on this page rather than solved. A worker is an OS process with the same filesystem permissions as the harness, so an allowed command can read any absolute path on the machine. No environment construction short of an OS-level sandbox prevents that. What the design does instead is remove every hint: no string in the staging surface names the real repository's location, so reaching the held-out surface would take a model inventing a path it has never seen. The raw harness log is archived per run and scanned at grading for the attempt.