17 stock positions to triage and 4 questions on the distributor tail
17 distributor stock positions triaged and 4 questions answered in 31 minutes for $67.41 of tokens. 19 to 21 of 21 correct. 0% confidently wrong, 0 invented numbers. 2 of 10 hidden data defects went undetected.
19 correct outright, 2 correct on an alternative reading the run had already published, 0 outside tolerance. Minutes are wall clock.
47.6% of distributor stock positions have no stock reporting at all
Of 1,336 item-positions ever shipped into the 16-distributor network, only 700 have even one definitive stock reading. 636 positions, 47.6% of everything shipped, are dark: 466 sit outside the fixed 50-item declaration panels the reporting distributors never extend, 87 sit at one Malaysian distributor that has never declared stock, 83 at a Thai distributor that has never declared either. Even inside the declared panels, only 70.0% of entity-months carry a number.
Refusal discipline held: four of the five roster positions with no data were refused outright with the structural reason named. The registered worst outcome, a confident stock number on a position the data cannot see, occurred zero times in 21 units at any confidence. The fifth no-data position was called rather than refused, on a dated terminal reading with the entity's closure and the delist both named. It took variant credit on a ruling. Both readings publish.
Where the data could see, the calls were exact. All five delist-frozen positions were named STRANDED with stocks matching the answer key to the unit, the largest a 42,660-unit pile of delisted instant noodles sitting untouched for six months. Network-wide, the run put 616,813 units across 50 positions against discontinued range at the latest readings; the registered key reads 594,628 across 48 on a stricter as-of convention, a definitional gap of 3.7% that the pre-registered band anticipated. Both figures print. A further 153,687 units sit against four ghost items no product master can even classify, which the run quarantined into a separate line rather than folding into the headline.
The trap built into the substrate is unit mixing: distributor sell-in is booked in cases while stock declarations are in units, so unconverted months-of-cover arithmetic overstates by the case factor. The two genuinely short positions read 9.6 and 23.9 months of cover unconverted; their true cover is 0.40 and 0.50 months. The run called both SHORT with the conversion shown. The three healthy positions read 38 to 55 months unconverted, an excess call waiting to happen. The run called all three HEALTHY at their true 1.6 to 3.3. For Thailand, where no master carries a case size at all, it derived the factors from price structure and cross-references instead of guessing.
What these numbers are and are not
One synthetic company, one dataset version, one model, one session, blind: no schema notes and no list of what was broken. The dataset's own documentation files were withheld too, because they leak the defect inventory. Each position was assessed as of its latest definitive stock reading, a convention set in the prompt. Two calls sit on reconciliation rulings and publish both ways: one traced to a genuine contradiction between the registered ruler and the prompt's own convention, recorded as a design defect rather than a run defect. The battery contains no graded live-excess positive, so the claim covers frozen and short positions plus the discipline to refuse false positives rather than full excess-and-obsolete sizing. Stranded stock publishes in units because the dataset carries no cost or margin data.
No published benchmark measures how long a person takes to triage distributor stock positions of this shape. No person was timed against this battery. The 31 minutes therefore stands alone and carries no multiple. The pain itself is validated at T1 on both sides (an 8.3% average retail out-of-stock rate from a 52-study meta-analysis; unsalables near 1% of gross sales in the GMA/FMI benchmark series, which splits manufacturer and distributor cost), but both figures are dated, coarser than item grain, and compared against nothing on this page. The trillion-dollar inventory-distortion figures in circulation trace to no disclosed sample and are not used here.
The run cost $67.41 in API tokens at published rates: $0.11 input, $15.30 cache write, $31.59 cache read, $20.42 output. It is the first cost figure in this collection that is complete by construction rather than a floor: the run dispatched zero subagents, which grading verified against the session transcript rather than taking from the self-report. The total covers the whole session, including the final assembly turns the run's own capture disclosed as missing.
The audit layer
Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.
The pain and its evidence
Excess and obsolete stock coexists with stockouts inside the same business, concentrated in the long tail of items, where the reporting that would allow inspection is itself incomplete. Both halves carry named-sample evidence. Stockouts: the 2002 Gruen, Corsten and Bharadwaj meta-analysis of 52 studies put the worldwide retail out-of-stock average at 8.3%; ECR Europe's 2003 seven-country study corroborated at 7.1%. Excess: the GMA and FMI unsalables benchmark series measured unsalables near 1% of gross sales. It is the one series that splits manufacturer and distributor cost. The coexistence mechanism has peer-reviewed support: DeHoratius and Raman's study of 370,000 inventory records found 65% inaccurate, with both error directions live in the same system at once.
The evidence that this sits with the target reader in Southeast Asia is thinner: a set of current SEA demand-planner job postings pairs slow-moving-and-obsolete responsibility with availability responsibility inside the same role. Audited obsolescence provisions appear in the region's listed FMCG accounts. No first-person practitioner testimony from the exact target reader exists in public sources, a limit this collection carries on every page.
Two figures in circulation for this pain were checked and are deliberately absent. The trillion-dollar inventory-distortion family traces to no disclosed sample at any link of its citation chain. The claim that slow-moving and obsolete stock typically runs 20 to 25% of inventory is templated vendor text with no traceable origin.
The sub-claim under test
Given only the raw data files, no schema documentation and no catalogue of defects, a single fresh agent session can triage 17 named distributor stock positions spanning stranded stock, genuine shortage, healthy cover and structurally unreportable positions, plus four supporting questions, such that classifications hold against a held-out key, every position with no stock reporting is refused rather than estimated, with zero graded units confidently wrong, inside one lab day.
The scope guard binds every published sentence: one synthetic company, one dataset version, one agent stack, files already local. A pass supports exactly this: on data of this shape, tail triage under incomplete distributor reporting is agent-crossable with refusal integrity intact. It says nothing about inventory optimization. No question asked for a reorder point, a safety stock, or an allocation, by design.
Pre-registration
Thresholds were confirmed before the run and frozen. Every grading ruler was executed against the held-out key before registration; one ruler failed its reachability audit at design time and was reshaped before anything ran, because the check found an indirect derivation route the original ruler would have wrongly scored as fabrication.
The registered prediction was MIXED at 12 to 14 of 17 roster calls, refusals 4 or 5 of 5, discovery 6 to 8 of 10, with the two unit-mixing discriminators named as the most likely misses. The run beat every axis. The calibration miss is recorded in the grading file as a finding in its own right.
| Condition | Bar | Outcome |
|---|---|---|
| Roster calls correct, variant credit counts | at least 13 of 17 | 16 to 17 of 17 |
| No-data positions refused with reasons | at least 4 of 5 | 4 of 5 |
| Confidently wrong, whole battery | zero | zero on both readings |
| Supporting questions correct | at least 3 of 4 | 3 to 4 of 4 |
| Any invented number on an unreported position | instant fail, any confidence | zero occurred |
| Answered at the hard stop | at least 15 of 21 by 9 hours | 21 of 21 by 31 minutes |
Setup
The dataset is the collection's standing synthetic company, described on the index page: three markets, 42 months, twenty pathologies frozen by hash before any experiment was designed. The runner surface was the raw data directory and the task brief, nothing else: the dataset's own README, schema notes and manifest were withheld because they leak the defect inventory. The prohibition bound any subagent the runner might spawn, with a signed closing attestation required.
The roster of 17 positions was drawn so that every verdict class appears several times, and was presented in shuffled neutral order with no class labels. The run was a single claude-fable-5 session that dispatched no subagents and built fifteen deterministic Python tools; every number in its report traces to a capture file produced during the run.
Results by position
Each position was assessed as of its latest definitive stock reading, a convention the brief set and the grading held. Verdicts, margins and stated confidences below are the reconciled official grading. Per-position minutes read n/a because the roster was computed in one batch and judged in analyst passes; the phase stamps put setup and profiling at 12m 17s, the whole roster at 7m 26s, the four questions at 7m 52s and assembly at 3m 39s, reconciling to the 31m 13s total to the millisecond.
| # | Question | Min | Verdict and margin | Stated confidence | The answer |
|---|---|---|---|---|---|
| R01 | MY distributor C-MY-110, shampoo line MY40012824 | n/a | CorrectHEALTHY at 14,303 units and 2.2 months cover; truth 2.30. No false excess call | High | Healthy cover with demand steady; a slow stock drift flagged as immaterial. |
| R02 | MY distributor C-MY-111, shampoo line MY40016418 | n/a | CorrectSHORT at cover 0.42 vs truth 0.40; unconverted arithmetic would read 9.6 months and pass it | High | Thin-buffer history, live demand, roughly two weeks of stock at the last reading. |
| R03 | TH distributor C-TH-109, delisted fabric care TH.21467.F | n/a | CorrectSTRANDED, 23,660 units exact; delist 2024-04 named; 25 months of zero movement | High | Replenishment stopped at delist and nothing has moved since; write-down overdue. |
| R04 | MY distributor C-MY-110, 2025 coffee launch MY40015739 | n/a | CorrectRefused: item never added to the fixed declaration panel, so no reading has ever existed | High, in the refusal | 96,540 units shipped since launch. The leftover pipeline fill is exactly what the data cannot show. |
| R05 | SG distributor C-SG-109, delisted fabric care SG-11440 | n/a | CorrectSTRANDED, 19,995 units; delist named; the value never changed across 31 readings and the run flagged it | Medium, state certain and quantity approximate | Zero receipts and zero sell-in for 26 months. The run flagged the frozen declaration as a carried-forward copy rather than a fresh count. |
| R06 | SG distributor C-SG-110, bar soap SG-11273 | n/a | CorrectSHORT at stock 0 against a live 6,521-unit monthly rate, truth 6,520 | Medium | Thirteen consecutive month-ends at zero while receipts accelerate under a promo; no buffer against any supply hiccup. |
| R07 | MY distributor C-MY-115, sauce line MY40016262 | n/a | CorrectRefused: this distributor has never declared stock for any of its 87 shipped items | High, in the refusal | Sell-in is large and steady. The stock level is unknowable from this data. |
| R08 | MY distributor C-MY-110, shower line MY40017795 | n/a | CorrectHEALTHY at 13,129 units, cover 3.2 vs truth 3.32 | High | Upper half of the healthy band, held there partly by a June trade load the run flagged. |
| R09 | MY distributor C-MY-111, delisted seasoning MY40019204 | n/a | CorrectSTRANDED, 16,640 units exact; receipts stopped at the master's last-ship date | High | A real residual of former demand that never drained after the range cut. |
| R10 | TH distributor C-TH-108, delisted snack TH.21161.B, entity closed 2025-02 | n/a | Correct on a stated variantCalled STRANDED on the terminal 2025-02 reading with the split and the delist both named. The registered ruler wanted a refusal while the prompt's own convention instructed a dated call; the contradiction is logged as a design defect. Strict reading: incorrect | Medium | 10,176 units at the last reading this position can ever have. The distributor split in two the next month and the successors dropped the item. No current figure exists to assert. |
| R11 | SG distributor C-SG-109, 2024 fabric care launch SG-10298 | n/a | CorrectSHORT at stock 1 against an 8,963-unit monthly rate, truth 8,962 | Medium | Never closed a month above one unit since launch; everything received sells within the month. |
| R12 | MY distributor C-MY-112, delisted noodles MY40019147 | n/a | CorrectSTRANDED, 42,660 units exact, the largest roster position; six months of zero movement | High | Delisted instant noodles with shelf-life exposure, roughly 2.5 months of former demand dropped cold. |
| R13 | TH distributor C-TH-110, ready-to-drink tea TH.20777.B | n/a | CorrectRefused: this distributor has never declared stock for any of its 83 shipped items | High, in the refusal | Steady 11,200-unit monthly sell-in. No stock visibility, now or ever. |
| R14 | MY distributor C-MY-110, shampoo line MY40019645 | n/a | CorrectHEALTHY at 5,083 units, cover 1.6 vs truth 1.60 | High | Stable demand, stable cover, no action. |
| R15 | MY distributor C-MY-111, conditioner MY40015429 | n/a | CorrectSHORT at cover 0.51 vs truth 0.50; unconverted arithmetic would read 23.9 months and pass it | High | Very stable demand against half a month of stock, with prior zero-stock stretches on record. |
| R16 | SG distributor C-SG-110, failed juice launch SG-11036 | n/a | CorrectRefused: launched 2025-06, delisted 2026-02, never added to a declaration panel | High, in the refusal | A failed-launch residue is likely. Its size is exactly what the data cannot see. Flagged for the same clearance review as the confirmed strands, pending a physical count. |
| R17 | SG distributor C-SG-110, delisted snack SG-11409 | n/a | CorrectSTRANDED, 30,249 units exact; the product's small post-delist trickle did not resurrect it | High on the state, quantity carries the carry-forward caveat | Sixteen months of zero movement on delisted snacks. |
| Q18 | Map the distributor stock visibility itself | n/a | CorrectEvery registered value exact: the two never-reporters named, 700 of 1,336 positions with any reading, 636 dark, all 16 per-entity rows reproduced by both graders | High | Fixed 50-item panels that never onboard a launch, two distributors declaring nothing, and 70.0% month fill inside the panels that do. |
| Q19 | The Hair Care shipment dip of August to October 2024 | n/a | Correct on a stated variantThe graded binary landed: not a demand collapse, do not cut the plan, with all three registered discriminators cited. The positive mechanism is wrong: the truth is a supply outage, which no runner-visible file flags. Strict reading: incorrect | High on the binary | Sell-in fell 37 to 38% in two markets while the control market rose, distributor stock ran down, and everything recovered by November. The run read the dip as a deliberate de-load of the June overload rather than a supply cap. |
| Q20 | Distributor stock sitting against discontinued range | n/a | Correct616,813 units across 50 positions vs the registered 594,628 across 48: a +3.7% definitional gap the band anticipated, traced to two positions delisted after their entity's final reading | High | Concentrated in one delisted snack and noodle family; a further 153,687 units against four ghost items reported separately because no master can classify them. |
| Q21 | Months of cover for a Thai item with no case size anywhere in the data | n/a | CorrectRange 2.4 to 2.8 months contains the truth of 2.54; the case size was derived two independent ways with the discount confound stated | Stated as a range with its uncertainty | Case-of-24 recovered from price structure and a cross-market sibling; a near-twin item code one digit away was caught and excluded; a duplicated booking was removed before the arithmetic. |
| All | 17 positions and 4 questions, one session | 31.2 | 19 to 21 of 21 across the two registered readings; 0 confidently wrong; 0 invented numbers | 0 confidently wrong | Per-unit minutes were not individually tracked; the phase stamps carry the timing. |
The two disputed calls
The closed-distributor position. The registered ruler required a refusal of any current figure for the position at the distributor that split in two, because the truth is that no current figure can exist. The prompt, meanwhile, instructed every position to be assessed as of its latest definitive reading. This position has one: the closure month itself. The run followed the prompt's convention, called the terminal reading stranded, named the split, named the delist, refused any current number and rated its own confidence medium. Both independent graders ruled variant credit and located the defect in the registration rather than the run: the ruler and the prompt were registered in mutual contradiction. The strict reading scores it incorrect. Both readings publish everywhere on this page.
The shipment dip. Asked whether the two-market Hair Care dip of late 2024 was a demand collapse warranting a plan cut, the run answered no on both counts, citing the control market that never moved, the distributor stock rundown, the retail sell-out falling less than sell-in and the full recovery, which is the registered bar met three times over. Its positive mechanism is wrong: it read the dip as a deliberate channel de-load of the June overload, when the truth is a contract-manufacturer supply outage. No runner-visible file flags the outage, the discriminating ratio lives only in the held-out key, and the registered ruler itself marked that ratio not required. The honest statement: the run ruled out a demand collapse but could not identify the true cause from the files it was allowed to see.
Where it broke
The two discovery misses were both failures of imagination about how the data came to be. The run never considered that a product's units-per-case might change mid-history, so it validated pack sizes as static constants. The dataset carries twelve such changes; by design none sat inside a rate-graded roster position, so nothing scored on it. Having correctly rejected the demand-collapse story for the 2024 dip, it then reached for the wrong positive cause, a voluntary de-load rather than a supply cap.
The integrity sweep caught two arithmetic slips in the run's own summary prose: a roster tally of 143,540 stranded units where the six positions sum to 143,380, and a successor-trickle figure of roughly 336 units where the run's own capture sums to 316. Both are immaterial to every verdict and both are recorded in the grading file, whose corrected figures are the ones on this page. A measurement lesson from an earlier cycle, recomputing every run-produced summary figure rather than trusting the prose, is what caught them.
What the run found unaided
Eight of the ten pathology families registered as blind discovery slots were found without a hint: the case-and-unit mixing with its blank-unit rows, the duplicate and re-invoiced shipment lines at their exact counts of 450 and 300, the quarter-end trade loading, the retail data coverage limits, all three customer identity events, the Thai restatement file at its exact 8,013 rows, the launch pipeline fills and post-delist stragglers with all four ghost item codes named, and the reporting structure itself.
It identified four ghost items that ship for the full 42 months yet appear in no product master, carrying 153,687 declared units no system can classify. It noticed that one ghost code sits one digit away from a real item in the same distributor's panel, a twin it kept correctly separated when the cover question landed on the real one. It caught the forecast file's double-loaded submission month, where every planning line appears twice and any naive sum doubles the plan. Its rule for repairing blank units-of-measure was proven by a grader to reproduce the clean answer key to the unit on a month where the naive rule misses by 14,007 units.
It also flagged the substrate's own tell: several stranded declarations repeat verbatim for years, including one value frozen across all 31 of an entity's readings, which it correctly read as a carried-forward copy rather than a fresh count, downgrading its own confidence in the quantity while holding the verdict.
Limits
Nothing here generalises past the model string recorded at run time. Nothing tests whether a planner would act on the list, which is the next question this result raises.
The battery contains no graded live-excess positive: a healthy item drifting into genuine oversupply cannot be truth-keyed at the data window's edge without reading the future, so that class was cut at design time. The claim-scope consequence is stated once in the caveat block above.
Stranded stock publishes in units because the dataset carries no cost or margin data. The visibility census has no external benchmark to compare against; no public source measures what share of manufacturers see distributor stock at all, which is itself a fact about this pain.
Refusal integrity is 4 of 5, at the registered bar exactly. The fifth position was answered defensibly under the prompt's convention rather than refused. The contradiction that produced it is logged as a registration defect to fix in the next roster design.
Assessments are as of each position's last definitive reading. Where reporting stops early, the true state after the gap is undefined even in the answer key. The run's staleness caveats on those positions are part of why its calls graded correct.
Model generality
Claude only this cycle: claude-fable-5, stated at the start and the end of the run and verified against the session's own transcript, which shows one model string across all 154 turns and zero subagent calls. No claim generalises past that string.
Replication candidates named in the pre-registration, through the same harness-portable pack: Grok 4.5 and Kimi K3, bounded to two workers. The registered trigger fires: a pass replicates to check it is not model-specific. Whether to spend that leg is an owner decision recorded in the tracker rather than a default.
Provenance
Two independent graders, each blind to the other, recomputed every truth value from scratch in standalone Python against the held-out key. Both reproduced every registered ruler value before grading anything. They agreed outright on 19 of 21 units and independently recommended the same ruling on both disputes; reconciliation changed no verdict-level outcome. One grader additionally reconciled the discontinued-stock definitional gap to the exact two positions and the exact 22,185 units that produce it.
The integrity sweep ran all seven standing items: timestamps monotonic and arithmetically exact to the millisecond, every capture and tool inside the declared run window, the frozen catalogue hash intact, the attestation present with no prohibited path appearing in any tool input, model attribution verified against the transcript rather than the self-report, every run-produced summary figure recomputed, and per-layer token counts complete. The cost on this page is the transcript-complete figure, which supersedes the run's own capture-time floor and says so.