Organisational Intelligence
Experiment 05PASSRan 30 to 31 July 2026Model claude-fable-5

7 questions on whether the sales history reflects real demand

Seven questions on demand signal in 71 minutes for $83.06 of tokens. 83% accuracy on one run, 50% on an identical rerun. 12 of 12 contaminated months found in both.

71 minSeven questions, one session
5 and 3 of 6Two identical runs of the same brief, split by one call about March
$83.06API tokens, the run only
The history as foundForecast built on the contaminated sell-in
9.69%
A different forecasting methodSame contaminated history
10.02%
The cleaned historySame method as the first bar
6.58%
Pooled weighted error across twelve held-out cells, six months by two markets, January to June 2026, from the graded run. All three bars are forecasts that run delivered as CSVs, graded by recomputing the truth from a held-out answer key. The second execution of the same brief reached the same conclusion in the same direction and size, and its numbers are set out under model generality rather than drawn here. The incumbent forecast already in the data covers only three of the six months and reads 4.69% on those, a figure the runs' own audit shows is flattered.
12 of 12Contaminated market-months found blind, in both runs
9.7% to 6.6%Holdout error once the signal was cleaned
0.3 ptsError moved by changing the forecasting method instead
9 of 12Hidden data defects found unaided, same three missed both times
23% lowThe graded run's one wrong number: Thailand lost demand

6 correct outright, 0 correct on an alternative reading the run had already published, 1 outside tolerance. Minutes are wall clock.

What moved the forecast error

Two changes were tested against the same six held-out months. Cleaning the contaminated history and forecasting from it moved pooled error 3.11 points, from 9.69% to 6.58%. Changing the forecasting method and leaving the history alone moved it 0.33 points, and in the wrong direction, from 9.69% to 10.02%. Against realised sell-in rather than true demand the same comparison reads 2.43 points against 0.83.

Almost all of that gain sits in June 2026, which is the month the loading pattern teaches a forecast to expect a season that consumers never produce. Dirty-history error there is 32.30% against 9.40% cleaned. The same cleaning made January slightly worse, 12.36% against 9.01%. So the honest statement is narrow: cleaning buys its error reduction where the contamination lives, at a small cost elsewhere.

Stated tentatively, because one run at one grain cannot establish it: the method effect landed at 0.33 points against 1.67 in this lab's own dry run, partly because the run chose two methods that perform similarly, so the size of the gap is a property of that choice as much as of the finding. What survives that caveat is the direction and the fact that the cleaning effect matched its dry run closely. And the incumbent forecast already sitting in the file still beats every variant the run built on the near-term months it covers.

This is the part of the cycle that repeated. The second execution of the same brief, the one that would have failed the diagnosis bar, reached the same accuracy conclusion in the same direction and roughly the same size, concentrated in the same June and costing a little accuracy in the same January, and it reproduced the incumbent's 4.69% to the digit. Its own look-ahead replay came back clean, so its figures carry the same evidentiary standard as the graded run's; they are set out under model generality below.

What these numbers are and are not

One synthetic company, one dataset version, one model, both runs blind on raw files and questions with no schema notes and no list of what was broken. The graded run passed at 5 of 6 diagnosis questions. A second execution of the same brief by the same model, on the same night, would have failed at 3 of 6, and the whole gap traces to one judgment call about whether a real but sub-threshold March lift counted as loading. Read the pass and the fail as one result with a wide band, never the pass alone. The error reduction is also concentrated in the one month whose seasonality the contamination fabricates, cleaning cost a little accuracy in January, and the incumbent forecast still beats every variant either run built on the three months it covers, on a number flattered by a defect the runs themselves uncovered.

Error is measured against true demand across six held-out months in two markets: 9.7% forecasting from the history as found, 6.6% from the cleaned history, 10.0% from changing the forecasting method instead. Those are the graded run's figures. The 71-minute wall clock is not a clean single-session measurement, because a concurrent second execution overlapped the first question and the pause. The 39% forecast error usually quoted for consumer goods is a global figure measured at product, promotion and channel grain; this test runs at market and month grain, which is materially easier, and the two are not comparable. The APAC-native evidence for this pain is thin, resting on one peer-reviewed Sri Lanka bullwhip study and a set of SEA job postings.

The run cost $83.06 in API tokens at published rates. This is the only experiment where that figure is complete rather than a floor. It splits into $35.31 for the coordinating session and $47.75 for the subagent it dispatched, whose usage records survived. The dispatched layer cost more than the coordination did, which is why every other run in this collection is published as a floor.

The audit layer

Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.

The pain

Consumer goods forecast error is widely quoted at about 39% against a 15 to 25% target, and the planner carries two conflicting targets at once, service level and write-offs, with no slack when error or bullwhip hits.

This experiment does not test that 39%. The figure describes real businesses at product, promotion and channel grain. This test runs at market and month grain, a materially easier environment, so no sentence here compares the two. The standing limit on this pain is printed rather than buried: the APAC-native evidence is thin, and the 39% is a global number.

BaselineValueSource and evidence tier
CPG average forecast errorAbout 39% against a 15 to 25% targetvaluechainplanning.com survey summary, sample undisclosed. Tier 3, global, not APAC
Best-in-class promotion forecast accuracy against laggards72% against 42%POI via TELUS. Tier 3 to 4
Upstream demand amplification in FMCGAbove 2 timesPeer-reviewed Sri Lanka FMCG study. Tier 1 academic, APAC-adjacent
Value of reducing error1 point of error is about 3% of pre-tax profitabilityIBF-cited, primary paywalled. Tier 1 lineage
SEA relevanceUnilever and Mondelez SEA job postings charter exactly this problemJob postings. Tier 3, APAC
The honest comparator at this grain4.69% on the three months it coversThe incumbent forecast inside this dataset, recomputed by both graders. Not an industry figure
What was tested

The claim under test, registered before the run: given only the raw data files, blind, under a hard information cutoff, one fresh agent session can detect that the distributor sell-in history is contaminated by quarter-end loading and name the affected months and markets without false positives, size the loading and its payback, separate a supply outage from loading, from promotions and from a genuine demand decline, split loaded-month spikes into loading and promotional uplift, rebuild a clean monthly demand series and deliver it as data, audit the incumbent forecast file, and show on a held-out six-month window whether cleaning the signal moves error more than changing the method does.

Blind means the runner received the seventeen data files and the questions. Held out from it: the answer key, the messiness catalogue, the schema documentation, the build log, the validation logs, the pre-registration, and every other experiment directory. Finding the mess unaided is the thing being measured.

Scope guard, binding on every sentence here: one synthetic company, one dataset version, one agent stack, one session, files already local, market and month grain. This does not test real demand-sensing systems, product-grain forecasting, real point-of-sale latency, or any model beyond the one named at the top of this page. A pass supports exactly one statement: on data of this shape, the demand-signal diagnosis layer of forecasting is agent-crossable blind, and cleaning the signal moved held-out error more than method choice did. It supports nothing wider.

The bar

Five pass conditions and four fail triggers, confirmed by the owner on 30 July 2026 as drafted, before anything ran. Every numeric tolerance traces to a grading script written and executed against the held-out key before the thresholds froze. No threshold moved afterwards.

The registered prediction was a pass at 5 of 6 with the accuracy claim supported, and it named the decomposition question as the most likely outright miss. The run landed at exactly 5 of 6 with the claim supported. The decomposition question survived, missing its predicted weak cell by 26.70 points but staying inside the allowance, and the attribution question fell instead.

ConditionRequiredActual
Pass: diagnosis questions correct5 or more of 6, including the reconstruction5, including the reconstruction
Pass: confidently wrong on the attribution question00
Pass: assumptions stated6 or more of 77
Pass: cleaning does no harm, and the leak check is cleanCleaned error at or below dirty error6.58% against 9.69%; leak check clean
Pass: wall clock with all seven answered8 hours or less71 minutes net, 7 of 7
Fail trigger: diagnosis questions correct3 or fewer5
Fail trigger: confidently wrong on the attribution questionanynone
Fail trigger: cleaning worsens the forecastmore than 1 point worse3.11 points better
Fail trigger: answered at the hard stopfewer than 5 of 77 of 7
Accuracy claim, subordinate and non-gatingCleaning effect at or above the method effect3.11 points against 0.33

Results

This record is the graded run. The second execution of the same brief is set out under model generality, question by question, with its own verdict.

Elapsed is wall clock as the run logged it. A further 5 minutes went on shared tooling before the first question. Confidence is what the run logged at the time, before grading.

The official score is 5 of 6 on the diagnosis questions, including the reconstruction, which the pre-registration made compulsory to any pass. Q07 is the accuracy leg; it scores as a claim that either holds or does not, and it is counted separately from the diagnosis total shown in the row below.

One timing deviation is on the record. The second execution ran concurrently on the same machine and overlapped this run's pre-tooling, its first question and its pause, with no overlap from Q02 onward. So the 14 minutes on Q01 and the 71-minute total are not clean single-session measurements. The box was 8 hours and the run finished 6.8 times inside it, so nothing in the verdict rests on the timing.

#QuestionMinVerdict and marginStated confidenceThe answer
Q01Where the sell-in history misrepresents demand, and by what mechanism14.0Correct12 of 12 loaded market-months named; mechanism, pattern and payback limbs all present; zero disqualifying false positivesHighQuarter-end and year-end distributor loading in Malaysia and Thailand, Singapore clean. One extra month was claimed as mildly loaded with no loading behind it, graded as a mechanism mislabel rather than an invention, and the run removed it itself later in the session.
Q02Size the loading and its payback6.0CorrectLoading totals Malaysia +5.76% and Thailand +14.54% against a 15% band; payback stated near 30% inside a registered 20 to 70 band, truth 45.27, direction correctHigh on the totals, medium on paybackBoth market totals inside band and the payback direction right. A secondary month-by-month limb, which does not score, came in at 7 of 10.
Q03The Hair Care collapse: attribution, window, damage, treatment10.0Outside toleranceAttribution, window, treatment and non-conflation all correct; Thailand lost demand 627,534 against a truth of 812,552, so 22.77% low against a 20% band. Malaysia passes at 6.40% lowHigh on diagnosis, medium on magnitudeRead as supply censoring rather than a demand decline, a promotion effect or channel loading, with the window and the treatment right. The named trap, taking the incumbent forecast as the counterfactual and concluding almost nothing was lost, was recognised and rejected.
Q04Split the loaded-month spikes into loading and promotional uplift5.0CorrectLoading dominant in 6 of 6 cells; the share is within 20 points on 5 of 6, with Malaysia June 2025 out by 26.70 pointsHighLoading, not promotion, drives every loaded month tested. The one cell outside tolerance is the exact cell this lab's own reachability work had flagged as the hardest.
Q05Rebuild the clean demand series and deliver it as data9.0CorrectShape error 3.815 Malaysia and 3.771 Thailand against ceilings of 5.00 and 5.50; contaminated-month improvement 9.92 and 13.85 points against floors of 5.4 and 3.5; clean-month no-harm inside allowance; level error 0.82% and 2.48% against a 20% bandMedium-highA 72-row monthly series delivered as a CSV that meets the delivery contract exactly, clearing every registered leg in both markets.
Q06Audit the incumbent forecast file9.0CorrectAll 4 registered defects found with evidence: missing submission months, a duplicated month, a reporting-grain switch, and the lag structureHighThe four registered defects, plus a fifth this lab had registered as not existing: the forecast file was built backwards from actuals that had not happened yet. Verified true by both graders.
Q07Forecast impact: cleaning the signal against changing the method4.0Accuracy leg supported, outside the diagnosis scoreBoth registered conditions hold. No-harm holds by 3.11 points; the cleaning effect of 3.11 points exceeds the method effect of 0.33, margin 2.78. The fail trigger, cleaning worse by more than 1 point, did not fireHighFour deterministic forecast files delivered and graded from a held-out truth. This is the accuracy leg and it scores as a claim that holds or does not, separately from the diagnosis count.
AllSeven questions, one session71.05 of 6 diagnosis correct, 1 outside tolerance. Accuracy leg supported, scored separately0 confidently wrongMedian 9.0 minutes per question.
The accuracy leg

Four caveats travel with every one of these numbers, and all four are on this page rather than in a footnote.

The margin is concentrated in June 2026, the month whose seasonality the contamination fabricates, and cleaning made January slightly worse at 12.36% against 9.01%. Cleaning buys its error reduction where the contamination lives, at a small cost elsewhere.

The method effect landed at 0.33 points against 1.67 in this lab's dry run, while the cleaning effect matched its dry run closely, 3.11 against 3.21. The margin is real and part of its size comes from the run choosing two methods that perform similarly.

The incumbent forecast wins near-term. On the three months it covers it reads 4.69% against 7.14, 6.98 and 6.59 for the run's three variants. And that 4.69% is flattered by the back-filling defect described below. Both facts print together.

Every absolute Thailand figure carries a shared unit-conversion bias, because the case-to-unit factor exists nowhere in the data the run was allowed to see. Only differences are quoted for Thailand, as the pre-registration required.

QuantityValue
Error on the history as found, 12 held-out cells9.69%
Error with a different forecasting method, same history10.02%
Error on the cleaned history, same method6.58%
Cleaning effect3.11 points
Method effect0.33 points
Same comparison measured against realised sell-in2.43 points against 0.83
June 2026, the month the contamination fabricates32.30% dirty against 9.40% cleaned
January 2026, where cleaning cost accuracy12.36% cleaned against 9.01% dirty
Incumbent forecast on the three months it covers4.69%, ahead of all three run variants on those cells
Delivery contracts on all five CSVsClean; the incumbent extract reproduces to the unit from an independent extraction
Where it broke

The attribution question, Thailand. Four of its five graded limbs came back right: the outage was read as supply censoring rather than a demand decline, the window was right, the treatment was right, and it was not conflated with the loading pattern next door. The fifth limb is the size of the damage. The run put Thailand's lost demand at 627,534 units against a truth of 812,552, so 22.77% low against a 20% band. Malaysia passes on the same limb at 6.40% low.

The failure is honest estimator error at the edge of the band. The pre-registration named a trap on this question, using the incumbent forecast as the counterfactual and concluding almost nothing was lost, and the run recognised that trap and rejected it explicitly. It carried the magnitude at medium confidence throughout. Its own alternative baseline would have landed inside the band; it led with the more conservative one and was graded on what it led with.

This particular miss did not repeat. The second execution of the same brief answered the same question inside band, 5.34% high in Malaysia and 0.73% low in Thailand, by building its counterfactual from a February-to-May 2024 mean where the graded run's estimator ran low. So the failure boundary here is an estimator choice rather than a capability wall, and the two runs disagree about which side of the band it lands on.

One further statement in the graded run's output is materially wrong and is deliberately absent from this page: a claimed uniform scope gap in the incumbent forecast file, which closes entirely once the brand-grain rows are counted. It is one of the four registered defects read backwards, both graders ruled it unsupported, and the registered scoring rule does not touch it. Both runs made the same mistake. It is recorded here rather than published as a finding.

Grading a run more kindly once you have seen its answers is how a lab stops being a lab. The strict letter cost the graded run this question, and the same strict letter is what let a threshold pass in its favour elsewhere.

What the run found unaided

Twelve hidden pathologies were registered as blind discovery slots before the run, scored but not gating. The graded run found 9, which is exactly the number the pre-registration predicted. The three it missed: a pipe-fill event hidden inside an already-loaded month, which the design predicted it would miss; the mechanism behind a reporting-grain switch, which it found as an artifact and then read backwards; and the re-invoiced half of the duplicate-document pathology, where it caught all 382 exact duplicates and never saw the 256 re-invoiced pairs. Two the design expected it to miss were caught instead. The second execution of the same brief scored 9 of 12 as well and missed the same three, which is the sharpest evidence in this cycle that the blind spots are a property of the model rather than of one session.

The finding that matters most went the other way, against the experiment itself. The audit question asked for four known defects in the incumbent forecast file, and the pre-registration asserted no other genuine defect existed there. The run found the four, then found a fifth: the forecast file had been built backwards from actuals that had not happened yet.

The evidence both graders verified independently. Forecasts submitted in mid-2024 for delivery three months later, months before a supply outage began, track the outage months to within a few percent while calling volume 31 to 43% below prior year, against a normal error near 11% at the same lag. No real forecaster predicts an unforecastable outage. The dataset's own pathology index confirms it by construction.

So the registered sweep is wrong and is logged as a deviation, credited to the run. It is also why the incumbent's 4.69% carries a caveat everywhere it appears on this page, and why the dataset's next version either generates that file from information available at submission time or documents the back-fit as intended.

Limits

One synthetic company, one dataset version, one model, files already local, market and month grain. Two executions exist rather than one, which is two points and not a variance estimate, and the second was never registered, so it is a windfall rather than a designed replication.

Those two points already bracket a pass and a fail. On a six-question battery with a 5-of-6 bar, one borderline detection call was worth three questions, so the formal verdict is more sensitive to a single judgment than the underlying capability is. Anyone reading the pass on its own is reading half the result, which is why the fail sits in the header.

The APAC-native evidence for this pain is thin: one peer-reviewed Sri Lanka bullwhip study and a set of SEA job postings. The 39% industry error figure is global and measured at a harder grain than this test, and nothing here is compared against it.

Thailand's case-to-unit conversion factor exists nowhere in the runner-visible data, so every absolute Thai number on this page carries the same conversion bias. Only differences are quoted.

The clean-month no-harm check in Thailand passes on the benchmark registered before the run, and on a benchmark rebuilt from the run's own better conversion the same check would breach its allowance. The registered letter governs, the same way it governed against the run on Thailand's damage estimate. It is logged as a threshold-construction fix for the next dataset version rather than as a run defect.

One pre-registered claim was falsified by the run. This lab had called Thailand's conversion gap unreachable blind and built the reconstruction thresholds on a weaker reference. The run inferred case sizes from price structure and beat that reference, so it cleared those thresholds with more room than the design intended.

Correct against a held-out key is not the same as a planner trusting the series. Nothing here tests whether a demand planner would plan against the reconstructed numbers, or what happens when the mess is unknown rather than merely undisclosed.

Model generality and run-to-run variance

Claude only this cycle: claude-fable-5, printed at the top of this page. No claim generalises past that string.

The cycle got an accidental replicate. A second execution of the same brief by the same model ran concurrently on the same machine, answered all seven questions in 24 minutes, and was found mid-flight by the graded run. No pre-registration provision covers it, so it is excluded from the official verdict. It was graded afterwards as a replication data point, by a single grader bound by the same reconciliation rulings, with lighter provenance than the official pass: the reconstruction and accuracy truths were recomputed independently, and the first four questions reuse the registered truth figures. Its look-ahead replay was run separately and came back clean, all five deliverables byte-identical, so its numbers below stand on the same evidence the graded run's do.

Under the same thresholds the replicate would have scored 3 of 6 and failed. The graded run scored 5 of 6 and passed. The whole gap is one judgment call at the detection stage: whether a real but sub-threshold March promotional lift counts as loading. The replicate headlined it as loading, which put three false positives into the first question, inflated Thailand's loading total to 27.64% against a 15 band (it reads 3.99% restricted to the six true loaded months), and displaced one cell of the decomposition question. Three questions on a six-question battery with a 5-of-6 bar. The graded run made the same call the other way, caveated it, and later retracted it itself.

What replicated is the part worth carrying. Both runs found all 12 true contaminated months. Both delivered near-identical reconstructions, shape error 3.822 and 3.580 against 3.815 and 3.771. Both reached the same accuracy conclusion, cleaning ahead of method choice, June-concentrated, January slightly hurt, with the replicate at 9.84% dirty, 11.27% on the method change and 7.00% cleaned. Both reproduced the incumbent's 4.69% to the digit. Both found exactly 9 of 12 hidden defects, missing the same three. Both read the same reporting-grain switch backwards, and both missed the same decomposition cell by the same 26.70 points.

The reading offered here, tentatively: the model-level blind spots replicated perfectly and the marginal judgment calls did not. A borderline attribution decision was worth three questions and flipped a formal verdict between two identical runs, which is a fact about how this battery scores as much as about the model. Anyone quoting the pass without the fail is quoting half a result.

Cross-model replication candidates named in the pre-registration, through the same pack: Grok 4.5, Kimi K3, Qwen3.7 Max, bounded to two workers.

Graded runReplicate, unregistered
Verdict under the same thresholdsPASS, 5 of 6FAIL, 3 of 6
Accuracy claimSupportedSupported
True contaminated months found12 of 1212 of 12
Disqualifying false positives03, all March, all headlined rather than caveated
Thailand loading total against a 15 band+14.54%+27.64%, or +3.99% restricted to the true loaded months
Thailand lost demand against a 20 band22.77% low, the graded miss0.73% low, inside band on a different counterfactual
Decomposition cells inside 20 points5 of 64 of 6
Reconstruction shape error, Malaysia and Thailand3.815 and 3.7713.822 and 3.580
Held-out forecast error, dirty and cleaned9.69% and 6.58%9.84% and 7.00%
Incumbent on its three covered months4.69%4.69%
Hidden defects found unaided9 of 129 of 12, the same three missed
Wall clock, all seven answered71 minutes net24 minutes
Provenance

Two independent graders, each blind to the other, each recomputing every truth value from scratch in standalone Python directly against the held-out key. Both reproduced every registered recipe value exactly before grading anything, and neither ran the run's own tooling to produce a truth value. Both reached the same verdict independently, and no reconciliation ruling changed a verdict-level outcome. The unregistered replicate was graded afterwards to a lighter standard, by a single grader bound by the same rulings, and its own look-ahead replay came back clean with all five deliverables byte-identical.

The look-ahead check was executed by a separate agent. All seven of the run's question tools were re-run unchanged against a copy of the data with every post-cutoff row stripped, and all five delivered CSVs regenerated byte-identical. A static audit found no prohibited-path reference in any tool or capture. Verdict clean, so the accuracy leg is valid.

Five deviations are on the record, none scoring. The run omitted its required closing attestation, and its substance was established independently by the replay, the static audit and file modification times. A second execution of the same brief overlapped the first question. The registered false-positive sweep on the forecast file was wrong, credited to the run. A pre-registered unreachability claim was falsified, also credited to the run. And a superseded design directory containing truth-derived values sat inside the experiment directory throughout the run, verified unread through zero references anywhere and modification times falling entirely outside the run window.

NextThe register