Organisational Intelligence
Experiment 09PASSRan 4 August 2026Model claude-fable-5

7 questions on why reported net sales moved

7 questions on why reported net sales moved, answered in 38 minutes for $65.12 of tokens. 21 to 26 of 29 checkpoints correct. 0% confidently wrong, 7 of 7 bridges closed. All 5 misses at the official grade are magnitude or population errors.

38.1 min7 questions and 7 bridges, one session, one model layer
21 to 26 of 2924 official across two blind graders; the strict and lenient readings publish as the range. 0 confidently wrong on every reading
$65.12API tokens, the run only, complete rather than a floor
7 of 7Bridges closed and internally consistent, six to the cent and the seventh to one cent
3 of 3Planted misattribution baits refused: the restatement as lost demand, the naive growth figure, the method artifact as trading
66.0MUSD of double-counted Thailand inside the quoted annual sales levels, rebuilt exactly
7 of 7Pure yes-or-no calls correct; every official miss is a magnitude or population error
11 of 12Hidden data pathologies in the graded window surfaced unprompted, against 8 to 10 predicted

2 correct outright, 1 correct on an alternative reading the run had already published, 4 outside tolerance. Minutes are wall clock.

At least 4 of the 7 quoted movements carried a data artifact of 12% or more of the printed number

Every number the questions quoted reproduced from the data to the cent or within its band, and then came apart. The half-year Thailand decline carries a net artifact of -17.8M THB, 21.3% of the movement, from a restated base and duplicated invoice lines. The group growth figure stands on a conversion-method artifact worth 19.3% of the growth it reports. The Thailand restatement question is 100% artifact: the pack number changed with no commercial event behind it at all. The Singapore May spike carries 12.0% of duplicated postings. These shares are grader-recomputed truth values, and separating them from the trading story is the task this cycle graded.

The largest single find is the growth number going upstairs. The quoted pair, 198.8M to 210.5M USD and growth of just under six percent, reproduces exactly as the naive sum of every row of the P&L file, a construction that counts seven restated Thailand months of each year twice. The run reproduced the naive construction, rejected it, and rebuilt the levels: 166,814,424.47 USD for 2024 and 176,454,954.03 for 2025, both exact to the cent against the held-out key, growth +5.78% as booked. The quoted levels carried 31.97M and 34.06M USD of double-counted Thailand, 66.0M across the two years, and the corrected bridge puts +3.6M to +5.5M of the +9.6M growth on currency depending on rate basis, with Malaysia driving the trading growth and Singapore declining.

Zero confidently-wrong held on every reading, and the mechanism is visible in the grading: the run declared its uncertainty exactly where it was weakest. All three planted misattribution baits were refused: the restatement was never sold as lost demand, the naive growth figure was never certified, the method artifact was never blamed on trading or currency. Six demoted magnitudes each carried a factor-2 confidently-wrong yardstick; none fired, because every wide estimate arrived labelled MEDIUM or carrying its caveat. Questions the run rated HIGH scored 88% of their checkpoints; questions it rated MEDIUM scored 75%.

The misses have one shape. Every pure yes-or-no call in the battery came out right, 7 of 7, and every bridge closed; the five official misses are all magnitude or population errors on components whose direction and cause were named correctly. The run also produced the most substantive attribution error just below the trigger line: it tagged the one genuinely growing Thai category, Home Care, as artifact in substance, a component worth 8.57% of its anchor, asserted at MEDIUM with the arithmetic disclosed. Scored wrong, fired nothing.

What these numbers are and are not

One synthetic company, one dataset version, one model, one session, blind: no schema notes and no list of what was broken. No budget or plan P&L exists in the dataset, so this tests period-over-period explanation rather than plan adherence. Accuracy publishes as a range because five checkpoints admit defensible stricter or more lenient readings: the official grade is 24, and the strict floor of 21 sits one point under the registered pass bar, so the strictest reading alone would not pass. Six loading and promotion magnitudes were demoted to reported-only before the run under an owner ruling: scored for direction, never size, each carrying a factor-2 confidently-wrong yardstick, none of which fired. A design session sharing the machine moved one of the run's tool files 81 seconds after the start; the collision touched no answer and is logged as a deviation, and the single-session framing carries that qualifier. The published 38.1 minutes corrects the run's own 37.4, which stopped at its last capture rather than the registered end stamp.

No person was timed against this battery and no published benchmark measures bridge-building on already-assembled data, so the 38.1 minutes stands alone and carries no multiple. The nearest published figure, APQC benchmarking data reported by CFO.com in 2021, puts the whole cycle to produce period-end management reports including variance commentary at a 12-day median with top performers at 7 days or fewer; it measures that whole cycle rather than this task and appears as pain context only. Every variance-specific magnitude in circulation traces to vendor marketing with no disclosed sample, and three widely repeated statistics for this pain proved on direct fetch to be fabrications with invented attribution; none is used here.

The run cost $65.12 in API tokens at published rates: $11.70 cache write, $29.13 cache read, $24.29 output, under a cent of uncached input. The figure is complete rather than a floor: the run dispatched zero subagents, verified against the session transcript rather than taken from the self-report, and the total includes the assembly turns after the run's own last capture. It was metered warm: three pre-run turns in the same harness session carried $1.26 of prior context, excluded and disclosed, and a cold standalone run would re-pay some of the cached reads at the 20-times-higher write rate.

The audit layer

Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.

The pain and its evidence

Variance commentary is chartered, manual and late. APQC Open Standards Benchmarking data, reported by CFO.com in 2021, measures cycle time to produce period-end management reports defined to include comments on variance from expected results: top performers 7 days or fewer after period end, median 12, bottom quartile 16 or more. Gartner, surveying 100 finance leaders in late 2023, found 66% naming the explanation of forecast and budget variances as the task where generative AI will have its most immediate impact, a direct signal the task is effortful today. AFP's 2025 FP&A Benchmarking Survey of 362 practitioners puts spreadsheets at 96% of planning and 93% of reporting, with data reliability and data accessibility as the top two obstacles.

Magnitude hygiene mattered more on this pain than any before it. Every figure in circulation that names variance analysis specifically, the days-per-analysis and share-of-analyst-time numbers, traces to vendor marketing with zero citations, including a widely repeated statistic attributed to a benchmark whose own pages do not contain it. Three more were caught as outright fabrications by direct fetch: a KPMG FX error-rate figure absent from the page cited for it, an NBER restatement study that does not exist in NBER's own index, and a method-shift figure whose cited source contains no number. The artifact share of reported movements, the quantity this cycle actually measures, is measured nowhere public at all.

The standing limit of this collection holds here for a sixth consecutive sweep: no first-person testimony from a commercial finance lead inside an APAC mid-sized FMCG was found in public sources. Regional job postings prove the duty is chartered in-region, and nothing else.

The sub-claim under test

Working blind on the frozen dataset, an agentic session can explain named period-over-period movements in reported net sales within one lab day: decomposing each movement into business causes and data artifacts, quantifying each named component within pre-registered tolerance, producing bridges that sum with an explicit residual, and never confidently attributing a data artifact to a business cause. Each of the seven questions is voiced as a GM would ask it, quoting the number as printed: the 4.28M to 3.66M November drop, the growth figure of just under six percent, the May spike of about twelve percent.

A budget-existence check ran before design and its finding binds the branch: no budget or plan P&L exists anywhere in the frozen dataset, the BUDGET entry in the FX file being a planning exchange rate rather than a plan. The experiment is therefore period-over-period explanation, month against month and year against year, and tests explanation rather than plan adherence; inventing a budget series was ruled out of bounds under any framing. This cycle owns backward-looking attribution only: what happens to the P&L if a lever moves is cycle 02, reconciliation pipelines are cycle 03, forecast-versus-actual is cycle 05, and model generality as a subject is cycle 06.

Pre-registration

Every grading ruler was executed against the held-out key before thresholds were set, so every band was proven reachable before anything was frozen. At the threshold gate the owner ruled the weakest estimators down rather than the bands wide: six loading and promotion magnitudes whose measured estimator spread ran from 93% of truth to wrong-signed were demoted to reported-only, their direction calls still scored as binaries, and the factor-2 boundary kept as a confidently-wrong yardstick. A wide band can no longer add a point; a confident hallucination past it still fails the run. The confidently-wrong trigger itself needs three legs at once: a miss beyond twice the band or a business-versus-artifact mistag on a component worth at least 15% of its anchor, confidence logged HIGH, and no caveat naming the uncertainty. The 15% leg was set because the measured method artifact in the group-growth question is 19.3% of its anchor; a 20% leg would have exempted the exact misattribution the cycle exists to catch. Three named instant-fire baits were registered: explaining the restatement as a trading decline, certifying the May spike's naive growth figure, and attributing the method artifact to trading or currency.

The registered prediction was MIXED to PASS at 21 to 25 of 29 with zero confidently-wrong and discovery of 8 to 10 hidden pathologies. The outcome landed at the top of the accuracy band and beat discovery on every reading. The three registered wobble candidates all passed; the five actual misses came from directions the prediction never named, which is recorded in the grading file as a calibration finding in its own right.

ConditionBarOutcome
Checkpoints correct, variant credit countsat least 22 of 2924 official; 21 to 26 across readings
Confidently wrong, whole battery including reported-only magnitudeszerozero, both graders, all yardsticks, all baits
Artifact flagships tagged and sized2 of 3, and 3 of 3 not confidently called business2 of 3 solid; 3 of 3 held
Bridges internally consistent7 of 77 of 7, six to the cent
Answered inside the timebox7 by the 8-hour box7 in 38.1 minutes
Setup

The dataset is the collection's standing synthetic company, described on the index page: three markets, 42 months, twenty pathologies frozen by hash before any experiment was designed. The runner surface was the task brief and the raw data directory, nothing else: the dataset's own README, schema notes and manifest were withheld because they name the defects column by column. The prohibition bound any subagent the runner might create, with a signed closing attestation required and delivered.

The run was a single claude-fable-5 session that dispatched zero subagents, built deterministic Python tools for every computation, and archived 20 capture files; every number in its report traces to a capture produced during the run. Each answer ships as a bridge table whose components carry business, artifact or mixed tags and an explicit residual line, which is the shape the grading scored.

Results by question

Verdicts, margins and confidences below are the reconciled official grading across two blind graders. The 38.1-minute wall clock decomposes on the run's own stamps: 0.4 minutes of access review, 6.0 of data profiling and tooling before the first question, 27.5 inside the seven questions, and 3.5 of assembly plus a 43-second closing tail. Median per question 3.5 minutes, longest 8.4 on the half-year bridge, zero paused minutes, largest gap between turns under three minutes. The run published 37.4 minutes measured to its last cost capture; the registered clock runs to the end stamp and the corrected figure is the one everywhere on this page.

#QuestionMinVerdict and marginStated confidenceThe answer
Q01Thailand, the half-year decline8.44 of 6 checkpointsAnchor -5.99% vs -6.35% truth, inside the band written for a dedup-imperfect current view; bridge closes exactly with the unexplained residual quantified; delisted-SKU units -7.95% off with all 8 delisted products found. The two misses: an artifact share summed to 25.1M THB gross where the truth nets to -17.8M, and a common-SKU price-volume-mix landing -40.6% from the total-portfolio truthMediumH1 2026 landed 62.4M THB under H1 2025. A June 2025 distributor load that never repeated, a Packaged Foods delist wave and broad volume decline on surviving SKUs carry it; promotion cushioned the fall. The run also projected the under-discounting pattern forward and flagged a live true-up risk of roughly 20M THB on the unrestated base.
Q02Thailand, where the decline lives2.24 of 4 checkpointsAll 3 graded categories inside tolerance with signs correct: -14.5%, -14.0%, -10.5%. The two category codes with no name in the lookup were crossed by cross-market majority vote, 40 of 40 SKUs right against a cross-reference that itself carries wrong links, pool -0.10% from truth; Packaged Foods share 60% vs 57.0%Medium, split stated per componentPackaged Foods is 60% of the decline: -26.6M THB from a six-SKU delist wave plus -13.4M of volume on survivors with zero price support. The near-miss recorded against this question: the one genuinely growing category, Home Care, was tagged artifact in substance.
Q03Singapore, the November cliff3.54 of 5 checkpointsAnchor -616,049.07 SGD exact to the cent; the commercial-definition alternative -11.38% vs -11.58%; the accrual swing 83,373.82 exact; bridge closes exactly. The miss: the answer handed 68% of the trading movement to seasonality where the truth baseline is flat and promotion is essentially the whole movementHighOctober's promo peak unwound into a November trough, and one accrual posted outside its promo window makes the printed cliff read worse than trading. The verdict delivered to the GM: do not treat November as a demand collapse.
Q04Group, the growth number going upstairs3.85 of 5, one on a stated variantCorrected annual levels exact to the cent with the naive construction reproduced and rejected; the Thailand double-count named; the method artifact sized -1.213pp vs -1.116pp truth, inside the 0.40pp band; the trading-versus-currency split correct on the clean rate basis with both bases declared, taking variant credit; the fiscal window named with the excluded 995,318.11 MYRHighThe quoted pair double-counts seven restated Thai months in each year. Corrected: 166.8M to 176.5M USD, +5.78% as booked, with +3.6M to +5.5M of the growth on currency depending on rate basis, Malaysia driving the trading growth and Singapore declining.
Q05Thailand, the number that changed behind us2.03 of 3 checkpointsExact to the cent on both prints, on the -23,327,288.13 THB delta, and on all seven restated monthly rows; the discount-only invariances shown on every row; the planted demand-decline reading refusedHighNobody changed a 2024 invoice: a retroactive distributor-discount true-up layer was posted September to October 2025 against 2024 months, proven by posting dates. Discount-only on every row, the volume story unchanged, no commercial action to take.
Q06Malaysia, the December that might not be demand4.51 of 2 checkpointsThe decision leg correct: the December premium went to distributor warehouses at 112% of the jump, every consumption proxy flat or down, the seasonal confound refused. The miss: the year-on-year anchor was built on all-channel value, +18.3% and -19.44%, where the checkpoint population is distributor units, +39.28% and -31.78%MediumDecember sell-in of +5.6M MYR went into six distributor warehouses, slightly less of it than last year, and January paid it all back. By December standards it was a soft December, and the run said so.
Q07Singapore, the May spike3.13 of 4 checkpointsGrowth +10.69% vs +10.65% truth; the re-invoiced phantom sized 84,815.29 vs 86,626 SGD with one 1,810.37 re-invoice missed; the naive 12% refused before anyone quotes it; bridge closes to one cent. The miss: the base-month volume direction was never statedHighA promotion wave at 3.7 times normal spend drove real volume and mix; 61,711.62 SGD of the jump is four postings booked twice. The defensible headline the run sent back: +10.7% invoiced, +9.4% on the P&L definition.
All7 questions, one session38.124 of 29 checkpoints official; 21 to 26 across the registered readings0 confidently wrongMedian 3.5 minutes per question.
Where it broke

Every one of the five official misses is a magnitude or population error on a component whose direction and cause were named correctly. The half-year artifact share was asserted as the sum of two oppositely signed artifacts, 25.1M THB gross, where the truth nets them to -17.8M; the run's own text implies the correct net and the checkpoint grades the assertion. The half-year price-volume-mix was built on common SKUs with delists broken out, landing -40.6% from the registered total-portfolio construction. The November-cliff answer named promotion first while its own decomposition handed 68% of the trading movement to seasonality against a flat truth baseline, and the checkpoint grades the decomposition. The December-loading anchor was computed on all-channel value where the checkpoint population is distributor units, a population swap that misses even on the run's own best distributor figure. And the May-spike answer never stated the base-month volume direction its checkpoint required.

Two findings the grading orders onto this page. The group-growth question sized the conversion-method artifact correctly, -1.213 percentage points against a truth of -1.116, without ever diagnosing what it was: the run called a six-decimal tie to the annual budget rate a set of conversion errors, and the artifact went upstairs sized but unexplained. And the near-miss: the one genuinely growing Thai category, Home Care, printed at +5.3M THB, was tagged artifact in substance at MEDIUM confidence with the arithmetic disclosed, 8.57% of its anchor, under the 15% trigger leg. The most substantive attribution error in the battery, scored wrong, firing nothing.

The integrity sweep failed its prose-accuracy item and the failure is recorded rather than smoothed: nine figures in the run's advisory prose are wrong, the material one stating a restatement component as -20.5M THB where the true figure is -23,903,050.55. Every bridge, anchor, cost and timing sum is exact, and none of the nine touches a scored checkpoint, which was verified item by item. The corrected figures are the ones on this page.

What the run found unaided

Eleven of the twelve hidden pathologies in the graded window were surfaced without a hint, against a registered prediction of 8 to 10; the strictest defensible reading is 9 and the softest 12, and all beat the prediction. The restatement layer was read correctly as a doc-level retroactive discount true-up, with all 3,554 restatement rows touching the compared base half traced to an existing invoice and the posting window proven from dates. The run also found the Thai months that appear twice in the P&L file and double-count any naive sum, Malaysia's 4-4-5 fiscal labels sitting beside calendar months in the same column, the 450 duplicated invoice lines, the three FX rate bases behind a conversion note that changes row by row, the blank units of measure, the mid-history distributor split, and the clashing date formats between files.

The strongest single discovery move: two Thai category codes have no name anywhere in the lookup, and the run recovered them as Packaged Foods and Beverages by majority vote across a cross-market reference that itself carries wrong links, correctly classifying all 40 affected SKUs. The one official miss is a pathology the run's promotion tagging presupposed but never named as a defect.

Limits

Nothing here generalises past the model string recorded at run time, and one run is one point: no variance estimate exists. The dataset is synthetic, built to the shape of a mid-sized APAC consumer business with its defects frozen before design.

No budget or plan P&L exists in the substrate, so nothing on this page tests budget-versus-actual variance, the phrasing most finance teams reach for first. The claim covers explaining movements between periods that both actually happened.

Six loading and promotion magnitudes are reported-only under the pre-run owner ruling: their directions were scored and their sizes were not, so no size accuracy is claimed for them anywhere. Their factor-2 hallucination yardsticks all held, which is a statement about declared uncertainty rather than about estimator precision.

The run shared its machine with a design session that, 81 seconds after run start, moved the runner's just-written profiling tool to a grader-side folder while executing an unrelated commit. The moved file is verified byte-identical with the runner's timestamp preserved, the contamination direction runs from runner to grader only, and no answer, stamp or scored item is affected; the deviation is logged in full in the grading file, and a working-tree freeze during live runs is registered as a protocol rule for the next cycle. The single-session claim on this page carries that qualifier.

The discovery denominator is the twelve pathologies with graded-window exposure, not all twenty in the catalogue. The artifact-share quantities this cycle measures have no external benchmark to compare against, because no public source measures them, which is itself a fact about this pain.

Model generality

Claude only this cycle: claude-fable-5 across all 145 run turns, verified against the session transcript's per-turn usage records rather than the harness's static system-prompt label, which disagreed and carries no usage record behind it. Three pre-run turns on a different model sat before the start stamp, excluded from the run and disclosed on their own cost line. No claim generalises past the verified string.

No replication leg was pre-registered for this cycle, so none is claimed or implied. Any multi-model run of this battery would be a separately declared addendum with its own registration, not an extension of this result.

Provenance

Two independent graders, each blind to the other, re-executed all seven frozen truth recipes with byte-identical outputs and independently recomputed every graded truth value from the raw data in standalone code, with zero mismatches on either side. An independent integrity auditor ran the seven standing sweep items plus twelve contract checks against the runner's session transcript, located and content-verified independently. One sweep item failed and is published above rather than smoothed.

The evidence chain is the strongest of any cycle audited so far: all 38 run artifacts carry file timestamps landing on their own log stamps to the second, the frozen dataset hash is intact with no run-day modification anywhere under it, the closing attestation is present character for character, the transcript shows zero prohibited reads, zero subagents across all 316 records, and zero network access. The $65.12 is the transcript-complete figure and supersedes the run's own capture-time floor of $64.30, which proved honest at 98 to 99.4% of actual on every count.

NextThe register