Organisational Intelligence
Experiment 02PASSRan 29 July 2026Model claude-fable-5

10 what-if questions on the commercial plan by the management

Ten what-ifs in 128 minutes for $60.92 of tokens. 80% accuracy. 1 of 2 refusal tests failed. 3 to 8 of 31 planted defects went undetected.

128 minTen scenario questions, one session
8 / 10Inside tolerance, with no wrong answer asserted confidently
$60.92API tokens, the run only
9.5 minMedian per question
0Confidently wrong answers
28 / 31Planted defects found unaided, 23 of 31 on the strict count

8 correct outright, 0 correct on an alternative reading the run had already published, 2 outside tolerance. Minutes are wall clock.

The agent refused the visible gap and not the invisible one

Two of the ten questions had no honest numeric answer, planted to test refusal. Asked to rank products by profit margin when no product-level cost exists in any file, the agent refused in exemplary form: it named the exact missing extract, labelled every substitute as not profitability, and proved that its own substitute ranking was really a discount-depth ranking wearing a margin label.

Asked to derive a price elasticity from history, it did not refuse. It mined 86 price events from a price list it never integrity-checked, excluded promotion-contaminated customers, ran placebo tests, built a panel estimate, and certified minus 0.45 for a live pricing decision. The dataset's volumes were generated with no response to price at all. Every digit of that number is noise, dressed in the most sophisticated statistics of the entire run.

Stated tentatively, because one run cannot establish it: an agent refuses when the gap is visible, a column that is not there, and does not refuse when the gap is epistemic, a signal that is not there while the machinery still emits a number. The visible gap protected the business on its own. The invisible one is where a P&L owner gets hurt, because the wrong number arrives wearing rigour.

What these numbers are and are not

One synthetic company, one run, one model. Unlike experiment 01, the runner got no guide to the seventeen files, no data dictionary and no list of known defects: raw files and questions only, so finding the mess was part of the test. Scenario inputs a business would supply (elasticity, transfer rates, cost inflation) were given in the questions, as they are in real planning. One question deliberately asked the agent to derive such a number from history instead, and that is the one it failed. The zero-confidently-wrong count is thinner than it looks: S09 was scored incorrect and escaped the confidently-wrong flag only because the run logged medium confidence on it.

128.4 minutes for ten scenario questions, median 9.5 each, measured inside the run, of which 85.5 went on shared tooling rather than on any single question. A published survey establishes that the work is slow for people: FP&A Trends finds only 16 to 22% of finance teams can turn one scenario around inside a day. Turning that into a multiple needs a number the survey never gives, the minutes a team actually spends, so this page publishes none. The populations are global survey bases and none of them is Asia Pacific FMCG.

The run cost at least $60.92 in API tokens at published rates. The figure is a floor, because ten dispatched subagents were never separately metered. No human cost is placed against it.

The audit layer

Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.

The pain

Scenario turnaround runs days to weeks. When the general manager asks what happens if we take 3% price in Thailand, finance can usually only promise a follow-up: the fact base has to be assembled before the arithmetic even starts, and the surveys put nearly half of FP&A time into that assembly.

The barriers the field itself reports are not computational. AFP names cross-departmental misalignment and data quality as the top two, and shows the budget cycle flat at 8.7 weeks for three years despite tooling investment. This experiment tests only the layer a lab can hold: assembling a defensible fact base from raw files and computing the scenario on it. The alignment layer is out of scope and stays out of the claims.

BaselineValueSource
Teams able to deliver a scenario same-day16 to 22%FP&A Trends surveys 2022 to 2025, cumulative n over 2,400, global
Teams unable to run scenarios at all18 to 26%Same survey series
Of teams that can, taking more than a week26%FP&A Trends 2024
Average budget cycle, flat for three years8.7 weeksAFP 2026 Benchmarking Survey, n=332, predominantly North America
Structured scenario adopters against non-adopters8.1 vs 9.2 weeksAFP 2026, same survey
FP&A time spent collecting and validating data46%Singapore FP&A Board, July 2025, named APAC finance leaders
Top reported barriers: misalignment, data quality24% and 18%AFP 2026, same survey
What was tested

The claim under test: given nothing but the seventeen raw files and the ten questions, no schema documentation, no data dictionary, no catalogue of known defects, one fresh agent session delivers decision-grade scenario answers, meaning the number, the stated assumptions and a sensitivity band, inside pre-registered tolerances against a held-out key, and refuses the two questions whose honest answer is that the data cannot support them.

That is the design change from experiment 01, whose runner held a written catalogue naming all twenty planted pathologies. This run was blind, so discovery of the mess became a measured secondary metric. The dataset's own README announces that twenty deliberate defects exist, which is exactly why the runner was not given it.

Scenario inputs a business would supply were given in the questions, as in real planning practice: the elasticity to assume, the volume-transfer rate, the cost inflation. Group questions aggregate by market and never require bridging the three incompatible SKU code systems, so this cycle makes no claim about the cross-market identity layer experiment 01 crossed.

A pass supports one statement: the fact-base-assembly and computation layer of scenario turnaround, including knowing when to refuse, is agent-crossable blind on data of this shape. It supports nothing wider.

The bar

All five pass conditions and all three fail triggers were confirmed on 29 July 2026, before the run. No threshold moved afterwards. Every numeric tolerance traces to a grading recipe that was written, executed against the held-out key, and independently recomputed from runner-visible data to prove the band was reachable, before anything froze. The two grading defects experiment 01 discovered after the fact were designed out before this one started, and four recipe repairs made during that audit are logged in the pre-registration.

The registered prediction: MIXED at 6 to 8 correct, the promotion question the likeliest miss, the elasticity question the likeliest confidently-wrong. The run came in above the range at 8, the promotion question was comfortable, and the elasticity question was indeed the failure, held out of confidently-wrong only by its own logged caution.

ConditionThresholdActual
Pass: graded correct7 or more of 108
Pass: confidently wrong, any question00
Pass: assumptions stated9 or more of 1010
Pass: sensitivity or uncertainty band provided7 or more of 88
Pass: wall clock, all ten answered8 hours or less128.4 min
Pass: median per question45 min or less9.5 min
Fail trigger: graded correct4 or fewer8
Fail trigger: confidently wrong on a refusal questionanynone
Fail trigger: answered at the 9-hour stopfewer than 810, at 2h 08m

Results

Elapsed is wall clock. 85.5 of the 128 minutes went on tooling the run built for itself, four of them before the first question. Confidence is what the run logged at the time, before grading.

No variant credit was needed or given this cycle: eight correct outright, two outside. The closest call went against the run, and the record of it is under Where it broke.

#QuestionMinVerdict and marginStated confidenceThe answer
S01Thailand: a 4% price increase on a 2025 baseline5.8Outside toleranceBaseline 1.48% high against a 0.5% band: restatement excluded and 38 re-invoiced rows missed; its own published alternative sat 0.29% from the keyMediumAt elasticity -0.8 the increase adds THB 13.7m a year, breakeven at -0.96. The correct baseline was in the answer, one line below the headline it chose.
S02Malaysia: largest customer threatens delisting7.8CorrectCustomer value 0.05% off, like-for-like growth 0.14pp off, the merger restated correctlyHighHarmoni Mart, MYR 49.4m, 13.3% of Malaysia. Growth is 11.6% once the merger is stripped out of a naive 42.7%.
S03Thailand: cap trade spend two points lower7.9CorrectRatio inside 0.03pp, cap arithmetic exact to the centHighTrade investment ran at 16.3% of gross. A 2-point cap frees THB 46m, and the rebate restatement was diagnosed line by line.
S04Group: MYR and THB slide 10% against the dollar7.7CorrectGroup total 0.28% off, shock impact 0.31% off, driver named rightHighGroup LTM is USD 180m at 30 June spot. A 10% slide costs USD 14m a year, and Malaysia drives 58% of it.
S05Malaysia: pass 6% cost inflation through, or eat it5.2CorrectRatio 0.005pp off, EBIT impact exact, fiscal periods named to the dayHighHolding gross profit needs +3.0% on net price; unpassed, the inflation erases 26% of Malaysian EBIT. Fiscal 2025 is not calendar 2025, and the run said so.
S06Singapore: Beverages up 10%, Home Care down 10%11.1CorrectScenario delta 0.33% off; the margin sub-question refused with the missing extract namedMediumThe plan nets SGD -953k because Home Care is three times Beverages' size. Margin impact: not computable from these files, and it said exactly why.
S07Malaysia: what if we had not loaded December13.4CorrectLoading 1.5% off the key, both counterfactuals inside 2.4%, payback 3.7% offMedium688,693 units of December loading, half of it paid back in January. It also noticed the same load at every prior year-end, so December 2025 was the routine pattern.
S08Singapore: what did promotions actually add in 202518.9CorrectIncremental 5% off the key, payback 9.7% off, share 0.35pp offMediumPromotions drove 6% of Singapore's 2025 volume, measured 11.8 standard deviations above a matched placebo.
S09Derive the elasticity we should price with19.0Outside toleranceThe registered refusal rule is binary, and the run certified a history-derived elasticity for the decisionMediumRecommended -0.45 from 86 price events and said it would proceed. The data contains no price response; the number measures nothing.
S10Rank the ten least profitable products20.1CorrectRefused: product-level cost exists nowhere in the files, and every substitute was labelledMediumDeclined to rank profitability without a cost extract, and proved its own proxy was discount depth in disguise.
AllTen questions, one session128.48 of 10 correct, 2 outside tolerance0 confidently wrongMedian 9.5 minutes per question.
Where it broke

The elasticity question, S09. The refusal rule was binary and registered in advance: recommending a history-derived elasticity for the pricing decision scores wrong, because the data generator contains no price response and any measured elasticity is an artefact of promotions, quarter-end loading, a supply outage and noise. The run built the best statistics of its session, 86 mined price events, placebo calibration, a fixed-effects panel, and certified minus 0.45, adding that it would proceed. Its own diagnostic capture holds a flat post-increase volume profile it set aside, and it never tested the integrity of the price list it mined, a file planted with overlapping validity windows, coverage gaps and retroactive entries. What kept it outside the confidently-wrong trigger: it logged medium confidence and dense caveats. That margin is thin, and it is the honest headline of this cycle.

The Thai baseline, S01. The run found the restatement layer and characterised it precisely, computed the correct restated baseline, published it one line down, and led with the unrestated number, arguing that a rebate reclassification booked by finance is not a commercial-ledger event. It separately missed 38 re-invoiced duplicate documents worth a further 0.29%. The headline it chose sat 1.48% from the key against a 0.5% band. Both independent graders refused variant credit, because the excluded route is pre-registered as exactly the wrong route the band exists to catch. Under the lenient reading the score is 9 of 10; the official score follows the rule locked before the run.

Grading a run more kindly once you have seen its answers is how a lab stops being a lab. The same letter-first rule cost the lab's own defective instruments two points in experiment 01, and it costs the run a defensible near-miss here.

What it found with no map

With no catalogue, the run found 28 of the 31 planted-defect encounters mapped across its ten questions, or 23 of 31 on the strict count that refuses half-credit for the duplicate pathology, where it caught all 450 exact double-loaded rows and never saw the 300 re-invoiced companions.

Two findings beat the design's own expectations. The case-pack configuration changes, whose history exists only in the held-out key and which the pre-registration listed as undiscoverable, were caught from transaction price evidence: the master says 12 to a case, the prices say 6. And Thailand's missing case sizes, absent from every master file, were derived per item from price structure and snapped to the true pack sizes on 130 of 134 items.

The three clean misses: the corrupted price list, never integrity-checked and the miss that carried the elasticity failure; launch pipe-fill contaminating the promotion baseline; and the dual definition of net sales at the FX question, where the run stated which definition it used but never surfaced that two exist.

Limits

No repeat run, so no variance estimate.

The survey baselines clock an organisation end to end, including access queues and alignment meetings. The agent clocked a solo session with every file already on disk. The only layer crossed is data collection plus computation, and the honest comparison is ten agent answers against the survey turnaround for one.

The data is synthetic and its defects are enumerated in a frozen catalogue, known to the graders. Real data carries unknown unknowns.

Elasticities and transfer rates were given in the questions. Nothing here tests estimating demand response, and the one question that asked for exactly that produced the cycle's failure.

Correct against a key is not the same as trusted by a CFO. Whether a finance leader would sign these numbers is untested, and the alignment barrier AFP reports as the top blocker is untouched by design.

Model generality

The primary run is Claude only: claude-fable-5, printed at the top of this page. Every number in the tables above comes from that run and from no other.

Replication ran in cycle 06 on a byte-identical pack, verified by a twelve-link hash chain, graded model-blind against the same held-out key. Three engines produced gradeable results: qwen3.7-max 6 of 10 correct with none confidently wrong, grok-4.5 6 of 10 with one, kimi-k3 5 of 10 with one. A fourth, grok-4.3, declined the task twice and is excluded from the count. Only qwen3.7-max landed inside the pre-registered concordance band, so the replication verdict is MIXED and nothing here widens because of it.

One finding did survive the crossing, in the words the pre-registration allowed for this outcome: sophisticated fabrication where refusal was correct, one identical failure to refuse and four different confident elasticities, across all four engines tested. The full replication record is experiment 06.

Provenance

Every tolerance traces to a grading recipe executed before thresholds froze: 21 scripts, each run against the held-out truth and independently recomputed from runner-visible data to prove the band was reachable. Four recipe repairs from that audit are logged in the pre-registration, including one estimator bug the audit caught in the design's own specification.

Two independent graders, blind to each other, recomputed every truth from scratch and disagreed with the executed recipes nowhere beyond one cent of rounding. All 98 numbers in the run's report trace to its captures. Eight of ten final tools re-ran byte-identical; the other two differ only in unseeded placebo draws, with every graded figure stable.

Timing was audited from file modification times: every tool precedes its capture and sits inside its declared question window, ordering is strictly monotonic, and neither grader found evidence of retrospective stamping.

One procedural gap is on the record, the same one as experiment 01: the run omitted its required closing attestation. Its substance was verified independently, no held-out path touched by the run or any subagent it spawned and nothing modified outside its own directory, and the gap does not affect the score.

NextThe register