10 questions on making the monthly numbers agree
Six unseen months reconciled on the first attempt, no code changes, for $3.43 of tokens. 10 of 10 checkpoints passed, 0 true breaks missed. 47 items still need a person.
10 correct outright, 0 correct on an alternative reading the run had already published, 0 outside tolerance. Minutes are wall clock.
What remained for a person to judge
Everything rule-shaped compounded into the asset. Two generations of duplicate document, 750 of them across both periods. A late Thai restatement of 8,013 rows applied exactly once because the pipeline is a pure function of immutable source files. Malaysia's 4-4-5 fiscal calendar rebuilt into calendar months from dated source. 404 product codes and 99 customer ids mapped, with three customer identity events canonicalised from evidence. After the build, all of that runs in 2 seconds of compute.
What remained for a human is commercial judgment. 47 items after the rerun, dominated by 42 trade-spend postings carrying no promotion id, about 7 a month, which no pipeline should allocate on its own. Behind them sit 5 item codes that invoice real revenue and exist in no product master, waiting on a master-data decision. Value at stake on the human items runs 1.8% of the clean base in Singapore, 3.1% in Malaysia and 3.3% in Thailand.
The trust boundary comes from the run itself. It named five things it would not run unattended: diagnosing a new class of break, allocating unallocated trade spend, absorbing a restatement file of a different shape, creating master data for unresolved codes, and signing off a month where the published Malaysian P&L covers only part of it. Its own tripwire is the zero-residual assertion, and it says to treat any non-zero cent as a stop-the-line event.
What these numbers are and are not
One synthetic company, one run, two models. The data is generated so that the correct rule ties out exactly, which is why six checkpoints landed at literally zero: those zeros prove the run found the right rule, and real ledgers never tie to the cent. The run was blind on the pathologies with one disclosed exception, recorded in the pre-registration before anything ran: the task brief told it trade spend belonged inside the exception queue's perimeter. The zero-confidently-wrong count is scoped to the three checkpoints registered for it rather than to all ten, with one factual error sitting outside that scope.
2.2 minutes to rerun all six held-out months in a single invocation, of which 2 seconds is the pipeline and the rest is hashing and evidence. The run projects a real monthly cycle at about 5 minutes of machine-plus-operator time plus roughly 8 queue items, and rates its own confidence in that projection MEDIUM because it extrapolates six months to a steady state. Building the asset took 79.8 minutes once. That figure hides concurrency, running through three subagents at roughly 2.4 concurrent agent-hours on a prepared dataset with every file already local, so it is not an implementation estimate for a real finance function. Both legs are measured inside the same run. Published evidence says the surrounding work is heavy for people: AFP/APQC 2019 (n=430+) splits FP&A time 25% analysis, 42% data gathering, 33% process admin, and FP&A Trends 2024 (n=2,400+) puts roughly 45% on data collection and validation. Neither covers reconciliation of this shape, and no person was ever timed doing it, so no human month is priced against the 2.2 minutes.
The rerun cost $3.43 in API tokens, metered across the nine agent turns inside the run's own logged phase-2 boundaries. That covered all six held-out months in one invocation, so it is not a per-month figure. The pipeline it invokes is deterministic Python, independently re-run by the integrity auditor and regenerating all ten output CSVs byte-identical, and 2 seconds of that 2.2 minutes is the pipeline itself. Building the asset cost the other $62.20 of the $65.63 session total, once. Both figures are floors, because three dispatched subagents were never separately metered. The $3.43 was also metered inside a warm session, where $2.32 of it is cache reads that a cold monthly run would partly re-pay as cache writes. No human cost is placed against any of it.
The audit layer
Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.
The pain
Finance spends more of the month making systems agree than analysing what they say. The target here is the recurring part: the same shipments-against-P&L-against-trade-spend argument reopened every period, with the same duplicates found again and the same definitional gap re-investigated. Sign-off workflow, live source systems and FX consolidation are out of scope and stay out of the claims.
The magnitude in circulation is wrong and worth correcting before any of this is quoted. The often-repeated claim that 80 percent of the time goes on data wrangling has no finance source at all: it traces to a 2016 CrowdFlower survey about data scientists, and it is contested even there. The defensible non-analysis share for FP&A is 65 to 75 percent, and that is the figure this page works from.
Every population below is a global survey base. One APAC corroboration exists and it is a single meeting note.
| Baseline | Value | Source |
|---|---|---|
| FP&A time split | 25% analysis, 42% data gathering, 33% process admin | AFP/APQC 2019, n=430+ finance practitioners, tier 2 |
| FP&A time on data collection and validation | roughly 45% | FP&A Trends 2024, n=2,400+, global, tier 1 to 2 |
| Same figure, APAC corroboration | 46% | FP&A Trends meeting note, Singapore FP&A Board 2025, tier 3, APAC |
| Defensible non-analysis share | 65 to 75% | The two surveys above, read together, per this lab's pain validation |
| Analyst-days a month on reconciliation of this shape | withdrawn | This lab's own assumption with no public source, withdrawn 1 August 2026 because nothing measured could be found underneath it |
What was tested
The claim under test: given only the raw extracts, with no schema documentation, no data dictionary and no pathology list of any kind, one fresh agent session builds a reconciliation asset over 36 months of messy three-market data, such that the clean base ties to the held-out key inside pre-registered bands, true breaks are surfaced in a register instead of netted into the numbers, documented structural differences are classified as structural rather than queued as errors, and the pipeline then reruns unchanged on the following six months at a small fraction of the build cost.
The split: build on 2023-01 to 2025-12, 108 market-month cells. Holdout rerun on 2026-01 to 2026-06, 18 cells, verified before registration to contain 112 true duplicate breaks across all 18 cells, all four phantom item codes still active, the definitional gap persisting, and one Malaysian fiscal period straddling the boundary. No dataset file was created or extended; the split is a read-time filter on a frozen substrate.
The task was a pipeline build rather than a question battery. Ten runner-facing questions each feed one or more of ten graded checkpoints. Analysis was restricted to build months until the phase gate, the pipeline was hash-frozen before the rerun, and a rerun needing code edits was pre-registered as a reusability failure rather than a do-over.
Experiment 01 already crossed one-off reconciliation correctness on this same dataset, and none of it is re-graded here. Its Q05 validated cross-market identity links at 100% effective precision and rejected all 21 truly wrong links with no good link discarded. Its Q08 reproduced the Thai restatement magnitude to a 0.0000% delta with the double-count detector never firing. Its Q09 surfaced both the commercial and the finance definition of net revenue with the bridge quantified rather than one number picked without saying so.
What experiment 03 adds, none of it touched in cycle 1: whether the reconciliation becomes a reusable asset that reruns on a later period with zero code changes; precision and recall of an exception register against the full population of true breaks rather than answers to point questions; the unflagged-absorption rate, which is the value share of true breaks netted into confident numbers without ever being surfaced; whether legitimate structural differences generate a false-alarm queue every period; and the build-against-rerun cost ratio, which is the capacity claim itself.
The bar
All six pass conditions and all five fail triggers were confirmed on 29 July 2026, before anything ran. No threshold moved afterwards. Unflagged absorption is the primary risk metric and the failure this experiment exists to catch, which is why it carries its own fail trigger at market-period grain rather than a pooled average.
Every band traces to a grading recipe that was written, executed against the held-out key, and adversarially audited before registration, with a measured reachability ceiling and measured failure-mode floors. Two recipe defects of exactly the class experiment 01 discovered after the fact were caught at design time. The first: the drafted 20% bridge tolerance discriminated nothing, because the residual it was calibrated on turned out to be 100% duplicate contamination read as accrual noise. The second: the drafted single-miss absorption clauses would have auto-failed a 99%-recall pipeline roughly 29% of the time, because 32 of the 112 holdout breaks individually exceed the 3% clause on their own.
The registered prediction was MIXED, with the page's headline expected to come from the bridge decoy or the queue composition. The bridge was called a coin flip by construction: the P&L carries a decoy trade-spend column running 4 to 5 times the ledger total at the opposite sign. The run detected it, measured it at 3.4 to 4.4 times, validated against the postings ledger instead, and landed at exactly zero. Discovery was predicted at 6 to 7 of 9 and came in at 9 of 9. The outcome was PASS on 10 of 10 checkpoints.
| Condition | Threshold | Actual |
|---|---|---|
| Pass: asset validity | CP-01 all three clauses, CP-06 both clauses | Both pass on both graders' numbers |
| Pass: break surfacing | CP-02 precision and recall 0.97, CP-03 at pass level in every market-period | 1.000 and 1.000; 0.000000% in all six |
| Pass: false-alarm discipline | CP-04 bridge and false-alarm clause, CP-05 method, no must-not class queued | All clear, decoy column avoided |
| Pass: reuse | CP-07 all clauses, CP-08 precision and recall 0.95 | Hashes identical on attempt 1; 1.000 and 1.000 |
| Pass: economics | Rerun 25% or less of build elapsed and 45 minutes or less, attempt 1 | 2.76%, 2.2 minutes, attempt 1 |
| Pass: honesty | Zero confidently wrong on R3, R6, R7; assumptions stated on 9 of 10 | 0 confidently wrong; 10 of 10 |
| Fail trigger: unflagged absorption | 20% of true-break value on build, 30% on holdout, or any of the 18 largest breaks missing | 0.000000%; 18 of 18 registered |
| Fail trigger: pipeline needed code edits to rerun | Any edit | None, attempt 1, hashes identical |
| Fail trigger: noise generation | False alarm on the bridge, a must-not class in the queue, or a queue over 500 items | None; 90 items |
| Fail trigger: confidently wrong on R3, R6 or R7 | Any | None |
| Fail trigger: questions answered at the 9-hour stop | Fewer than 8 of 10 | 10 of 10, at 1h 27m |
Results
Each question is graded through the checkpoints it feeds, and every checkpoint came back PASS on both independent graders' numbers. The per-question minutes are short because the pipeline already existed by the time the questions were answered: 68 of the 86.7 wall-clock minutes went on building tooling, and the ten answers took 15 minutes between them. R7 and R8 were answered inside one work block, so their minutes overlap rather than sum. Confidence is what the run logged at the time, before grading.
No variant credit was needed. The re-invoice disposition was pre-registered as a coin flip the data cannot resolve, with grading set to recompute truth under whatever convention the run stated. The run derived a keep-original rule and validated it externally against the P&L in every one of the 108 market-months, and it coincided with the key.
Six of the ten checkpoints landed at exactly zero. On this substrate that means the run found the correct rule, and nothing finer. The checkpoints carrying real information are the Malaysian method ruling, the mapping reachability ceiling, the queue composition, the economics, and the discovery gaps.
| Checkpoint | Measured | Band |
|---|---|---|
| CP-01 clean-base integrity | 0.0000000000% over 108 cells; 0.00 arithmetic; 14 of 14 Thai cells exact | 0.002% max cell delta, restatement once |
| CP-02 build break register | 1.000 precision, 1.000 recall, 638 true, 0 false, 0 missed | 0.97 and 0.97 at pair grain |
| CP-03 unflagged absorption | 0.000000% in all six market-periods; 18 of 18 largest breaks registered | 8% build, 15% holdout, per market |
| CP-04 Singapore and Thailand bridge | 0.000000% in all four window cells; three false-alarm limbs clear; decoy column avoided | 3% residual at window grain |
| CP-05 Malaysian fiscal re-base | Method passed by ruling; 5.1968% worst month; 0.3844% on the tripwire cell; June gap disclosed at 3 of 30 days | Method binary, 18% sanity, 30% tripwire |
| CP-06 mapping completeness and lineage | 0.0000% dropped; 404 of 404 codes mapped or queued; 4 of 4 phantoms surfaced; 3 of 3 lineage events canonical | 0.05% dropped value, phantoms surfaced |
| CP-07 zero-edit rerun and holdout tie-out | Hash lists byte-identical, attempt 1, 0.0000000000% over 18 cells | Identical hashes, 0.002%, no degradation |
| CP-08 holdout break register | 1.000 precision, 1.000 recall, 112 true, 0 false | 0.95 and 0.95 |
| CP-09 queue composition | 90 items, 47 needing a human, zero must-not classes, 36 of 36 unallocated postings present | 500-item ceiling, no must-not class |
| CP-10 capacity economics | 2.2 minutes against a 79.8-minute build, so 2.76%, attempt 1 | 25% of build and 45 minutes |
| # | Question | Min | Verdict and margin | Stated confidence | The answer |
|---|---|---|---|---|---|
| R1 | One clean net sales number per market-month | 1.0 | CorrectCP-01: max cell delta 0.0000000000% over 108 cells against a 0.002% band; arithmetic residual 0.00 in local currency; restatement applied exactly once | High | 108 market-months delivered, every adjustment itemised as its own labelled column and reversible. Clean shipments net minus the off-invoice ledger equals the P&L to the cent in all 108. |
| R2 | What double-counted value sits in the extract | 1.0 | CorrectCP-02: precision 1.000 and recall 1.000 at pair grain against a 0.97 bar, 638 true breaks, 0 false positives, 0 missed | High | SGD 1.16m, MYR 6.72m and THB 41.6m of double count in two generations, 382 byte-identical repeats and 256 re-keyed copies. The P&L adjudicated which copy was real. |
| R3 | Why Singapore and Thailand shipments disagree with the P&L | 3.0 | CorrectCP-04: bridge residual 0.000000% in all four window cells against a 3% band, all three false-alarm limbs clear, and the decoy P&L trade-spend column detected and avoided | High | Off-invoice trade spend plus duplicates plus the restatement account for the whole gap, residual 0.00 in 72 of 72 months. The monthly hunt can stop and be replaced by a one-line assertion. |
| R4 | Why Malaysia never ties to a calendar month | 1.0 | CorrectCP-05: method clause passed by reconciliation ruling; worst month 5.1968% against an 18% sanity band; 0.3844% against the 30% January tripwire; the June coverage gap disclosed and quantified at 3 of 30 days | High | Malaysia books on 4-4-5, and only 112 of 1,096 build-window days sit in a fiscal period that matches their calendar month. Calendar months were rebuilt from dated source, residual 0.00 in 36 of 36. |
| R5 | The value still unmapped | 1.0 | CorrectCP-06: 0.0000% of value dropped against a 0.05% ceiling; 404 of 404 codes mapped or queued; 4 of 4 phantom codes surfaced unresolved; 3 of 3 identity events canonical; zero closed-id leakage | High | 5 item codes unresolved with THB 85.4m and MYR 10.5m riding on them, 37 more resting on corroborated weak evidence, and every customer id resolved. Nothing was dropped and every identity is priced. |
| R6 | Why the Thai adjustment file applies exactly once | 2.0 | CorrectCP-01(c): 8,013 rows registered at -47,230,338.68 THB, equal to the file total, and matching the restated P&L to the cent in all 14 affected months; 14 of 14 Thai cells exact | High | A delta ledger of retro trade claims delivered up to 15 months late. Double application is structurally impossible because the pipeline never writes back into the source it reads. |
| R7 | The rerun on six months it had never seen | 2.0 | CorrectCP-07: tools hash list byte-identical before and after, attempt 1, holdout tie-out 0.0000000000% over 18 cells; CP-08: precision 1.000 and recall 1.000 against a 0.95 bar, 112 true breaks, 0 false positives | High | Ran unchanged on 2026-01 to 2026-06 in 2 seconds of compute, first attempt. Residual 0.00 out of sample in all 18 market-months, and the one edge case degraded exactly as declared. |
| R8 | What is left in the queue after the rerun | 2.0 | CorrectCP-09: 90 items against a 500-item blow-up ceiling, 47 needing a human, zero must-not classes, all 36 unallocated postings present | High | 47 items for a person, worth 1.8% to 3.3% of the clean base depending on market. 42 of them are trade-spend postings with no promotion id, the same rate as the build period. |
| R9 | Build cost against rerun cost | 1.0 | CorrectCP-10: rerun 2.2 minutes against a 79.8-minute build, so 2.76% against a 25% ceiling and a 45-minute ceiling, attempt 1 | High on the measured costs, medium on the forward projection | Build 79.8 minutes of wall clock, rerun 2.2 minutes end to end including its evidence trail. The projected marginal month is minutes of machine time plus a small human queue. |
| R10 | The master data fix list | 1.0 | CorrectDiscovery metric, non-scoring: 9 of 9 pathologies detected, 6 of 9 on the strict all-limb reading. One factual error inside the register itself, outside every band | High | 12 data quality issues with measured impact, 9 worked around by the pipeline and 3 that no pipeline can fix without inventing facts. Four of them surfaced because a reconciliation failed. |
| All | Ten questions, one session | 86.7 | 10 of 10 correct, 0 outside tolerance | 0 confidently wrong | Median 1.0 minutes per question. |
Where the boundary is
Nothing fired at the bands, so the honest material is at the edges of the design rather than in the score.
The Malaysian method clause, CP-05. Both graders found the same indeterminacy: the run neither allocated fiscal figures by day count, which was the clause's pass wording, nor mapped period labels without day weighting, which was its fail trigger. It rebuilt Malaysia's calendar months exactly from dated source data, verified independently by both graders at zero residual across all 42 published fiscal periods on gross, trade spend and net. The clause was scored on its stated purpose, which is defeating the naive period-label read, and the run beats that read by 5.6 times on worst monthly error and reads 0.3844% against its 9.69% on the tripwire cell. Ruled pass by intent, with the letter gap recorded as a design deviation. Both readings publish. The run had also implemented and measured day-weighted allocation as a fallback and rejected it on the evidence.
The mapping reachability ceiling, CP-06. Of 180 multi-code truth product groups, 27 are unreachable from the data the run was allowed to see: they carry zero cross-reference links and their attribute signatures collide. Of the 124 that are reachable, the run grouped 124 correctly with zero over-merges, at link precision 1.0000 and link recall 0.8114. So the run sat exactly at the measured ceiling. The equivalence-class criterion behind that diagnostic carried no pre-registered threshold, and a threshold cannot be invented after a run, so nothing was scored against it. This is the surviving design defect of the cycle-1 class: cross-reference coverage was audited for the four phantom codes and never at group level.
The discovery limbs. The official metric is pathology-level and reads 9 of 9. Grader B's strict all-limb reading is 6 of 9 and publishes alongside it. Three limbs were never surfaced or never named: a case-configuration change partway through the history, the promotion-to-posting lag, and the quarter-end true-up cadence. One more was handled without being quantified: the split ratio between the two successors of the closed Thai customer, where 54.1 against 45.9 is measurable from the data.
One factual error inside the run's own fix list. It recorded the 252 blank-promotion-id build postings as all typed unallocated, where the truth splits them 216 unallocated and 36 true-up, and its own capture from the bridge work held the correct split. That question feeds the discovery metric rather than a band, so it is outside the confidently-wrong scope and outside every tolerance. It publishes as a blemish.
One procedural breach, disclosed by the run. A subagent ran a single read-only git status through a shell wrapper, against a run rule that banned git commands outright. The subagent disclosed it in its own log, and the orchestrating session refused to sign the closing attestation over it and reported the breach line by line instead. The integrity audit confirmed zero repository state change and no other git activity attributable to the run. Non-scoring, and the refusal to sign is the wanted behaviour.
Limits
The substrate is exact. On this synthetic data the truth is an algebraic function of the source files under the correct rule, so a zero-delta tie-out discriminates "found the correct rule" and nothing beyond it. Real reconciliation never ties to the cent, and the claim being made is about behaviour, surfacing a break rather than absorbing it, never about the width of the band.
The Malaysian checkpoint is method-graded because no flat tolerance at any grain separates an honest day-weighted re-base, whose worst honest month runs 13.6%, from a lucky naive mapping, whose best month runs 0.5%. A right method with botched arithmetic inside the 18% sanity band would not have been caught.
No new product launched in the holdout window, so the rerun tested whether the mapping asset survives delist churn and persistent phantom codes. Onboarding a new item code was not tested at all.
The wall clock hides concurrency. The build ran through three subagents, roughly 2.4 concurrent agent-hours inside 79.8 wall minutes. Wall clock is the pre-registered definition of the economics, and the agent-hours are the honest footnote to the 79.8.
The APAC evidence for this pain is one meeting note, and no public baseline anywhere covers reconciliation of this shape. Until 1 August 2026 this page held the run against an assumed 1 to 3 analyst-days a month. That range came from this lab rather than from a source or a timed person, so it was withdrawn and the chart now shows only the run's own two legs.
One company, one dataset version, one run split across two models, files already local, local currency only. No repeat run, so no variance estimate.
Model generality
Claude only this cycle, across two models: the run opened on claude-opus-5 and switched to claude-fable-5[1m] 31 minutes in, both printed at the top of this page. No claim generalises past those two strings.
Replication was pre-registered as triggered by exactly this outcome. A pass sends the same pack, unchanged and harness-portable, through the multi-model worker harness bounded to two workers, to check that the capacity claim is not an artefact of one model family.
Provenance
Two independent graders, each blind to the other, each recomputing every truth value from scratch in standalone stdlib Python against the held-out key and the frozen bands. Neither executed any of the run's tooling. Their per-checkpoint numbers agree exactly everywhere. The three judgment items, the Malaysian method letter, the mapping recall criterion and the true-up labelling, were handed up and ruled on in the record with both readings published.
An independent integrity audit checked timestamps, file modification times, both hash lists, a prohibition sweep of every file-open in the pipeline, and the attestation, then re-ran the whole pipeline: both period invocations completed with no edits and regenerated all ten output CSVs byte-identical. Verdict clean, with two evidence-handling deviations recorded. One tool and its capture predate the declared phase boundary by about 97 seconds, a narrative inconsistency with no numeric effect. Two subagent captures show created and modified times differing, and no graded number rests on them.
Bands were frozen before the run and every one traces to an executed recipe. The pair-grain repair that came out of the design audit was load-bearing on this exact run: scored at document grain, the build register's precision would have read 0.0737 against a true 1.0000, because a compliant pipeline's 8,013 restatement adjustments would have counted as false positives.
Timing is self-reported under the prompt's anti-batching rules, with the file-modification audit as the control. It came back monotonic and arithmetically consistent within 0.42 minutes.