Organisational Intelligence
Experiment 08PASSRan 3 August 2026Model claude-fable-5

17 stock positions to triage and 4 questions on the distributor tail

17 distributor stock positions triaged and 4 questions answered in 31 minutes for $67.41 of tokens. 19 to 21 of 21 correct. 0% confidently wrong, 0 invented numbers. 2 of 10 hidden data defects went undetected.

31 min17 positions and 4 questions, one session, one model layer
19 to 21 of 21Two disputed calls graded both ways and published as the range. 0 confidently wrong on either reading
$67.41API tokens, the run only, complete rather than a floor
4 of 5No-data positions refused outright, each with its structural reason named
0Invented stock numbers across all 21 graded units, at any confidence
47.6%Of shipped distributor item-positions carry no stock reporting of any kind
616,813Units frozen against delisted products at the latest readings, 50 positions
8 of 10Hidden data defects found unaided; a mid-series pack-size change and a supply outage missed

47.6% of distributor stock positions have no stock reporting at all

Of 1,336 item-positions ever shipped into the 16-distributor network, only 700 have even one definitive stock reading. 636 positions, 47.6% of everything shipped, are dark: 466 sit outside the fixed 50-item declaration panels the reporting distributors never extend, 87 sit at one Malaysian distributor that has never declared stock, 83 at a Thai distributor that has never declared either. Even inside the declared panels, only 70.0% of entity-months carry a number.

Refusal discipline held: four of the five roster positions with no data were refused outright with the structural reason named. The registered worst outcome, a confident stock number on a position the data cannot see, occurred zero times in 21 units at any confidence. The fifth no-data position was called rather than refused, on a dated terminal reading with the entity's closure and the delist both named. It took variant credit on a ruling. Both readings publish.

Where the data could see, the calls were exact. All five delist-frozen positions were named STRANDED with stocks matching the answer key to the unit, the largest a 42,660-unit pile of delisted instant noodles sitting untouched for six months. Network-wide, the run put 616,813 units across 50 positions against discontinued range at the latest readings; the registered key reads 594,628 across 48 on a stricter as-of convention, a definitional gap of 3.7% that the pre-registered band anticipated. Both figures print. A further 153,687 units sit against four ghost items no product master can even classify, which the run quarantined into a separate line rather than folding into the headline.

The trap built into the substrate is unit mixing: distributor sell-in is booked in cases while stock declarations are in units, so unconverted months-of-cover arithmetic overstates by the case factor. The two genuinely short positions read 9.6 and 23.9 months of cover unconverted; their true cover is 0.40 and 0.50 months. The run called both SHORT with the conversion shown. The three healthy positions read 38 to 55 months unconverted, an excess call waiting to happen. The run called all three HEALTHY at their true 1.6 to 3.3. For Thailand, where no master carries a case size at all, it derived the factors from price structure and cross-references instead of guessing.

What these numbers are and are not

One synthetic company, one dataset version, one model, one session, blind: no schema notes and no list of what was broken. The dataset's own documentation files were withheld too, because they leak the defect inventory. Each position was assessed as of its latest definitive stock reading, a convention set in the prompt. Two calls sit on reconciliation rulings and publish both ways: one traced to a genuine contradiction between the registered ruler and the prompt's own convention, recorded as a design defect rather than a run defect. The battery contains no graded live-excess positive, so the claim covers frozen and short positions plus the discipline to refuse false positives rather than full excess-and-obsolete sizing. Stranded stock publishes in units because the dataset carries no cost or margin data.

No published benchmark measures how long a person takes to triage distributor stock positions of this shape. No person was timed against this battery. The 31 minutes therefore stands alone and carries no multiple. The pain itself is validated at T1 on both sides (an 8.3% average retail out-of-stock rate from a 52-study meta-analysis; unsalables near 1% of gross sales in the GMA/FMI benchmark series, which splits manufacturer and distributor cost), but both figures are dated, coarser than item grain, and compared against nothing on this page. The trillion-dollar inventory-distortion figures in circulation trace to no disclosed sample and are not used here.

The run cost $67.41 in API tokens at published rates: $0.11 input, $15.30 cache write, $31.59 cache read, $20.42 output. It is the first cost figure in this collection that is complete by construction rather than a floor: the run dispatched zero subagents, which grading verified against the session transcript rather than taking from the self-report. The total covers the whole session, including the final assembly turns the run's own capture disclosed as missing.

The audit layer

Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.

The pain and its evidence

Excess and obsolete stock coexists with stockouts inside the same business, concentrated in the long tail of items, where the reporting that would allow inspection is itself incomplete. Both halves carry named-sample evidence. Stockouts: the 2002 Gruen, Corsten and Bharadwaj meta-analysis of 52 studies put the worldwide retail out-of-stock average at 8.3%; ECR Europe's 2003 seven-country study corroborated at 7.1%. Excess: the GMA and FMI unsalables benchmark series measured unsalables near 1% of gross sales. It is the one series that splits manufacturer and distributor cost. The coexistence mechanism has peer-reviewed support: DeHoratius and Raman's study of 370,000 inventory records found 65% inaccurate, with both error directions live in the same system at once.

The evidence that this sits with the target reader in Southeast Asia is thinner: a set of current SEA demand-planner job postings pairs slow-moving-and-obsolete responsibility with availability responsibility inside the same role. Audited obsolescence provisions appear in the region's listed FMCG accounts. No first-person practitioner testimony from the exact target reader exists in public sources, a limit this collection carries on every page.

Two figures in circulation for this pain were checked and are deliberately absent. The trillion-dollar inventory-distortion family traces to no disclosed sample at any link of its citation chain. The claim that slow-moving and obsolete stock typically runs 20 to 25% of inventory is templated vendor text with no traceable origin.

The sub-claim under test

Given only the raw data files, no schema documentation and no catalogue of defects, a single fresh agent session can triage 17 named distributor stock positions spanning stranded stock, genuine shortage, healthy cover and structurally unreportable positions, plus four supporting questions, such that classifications hold against a held-out key, every position with no stock reporting is refused rather than estimated, with zero graded units confidently wrong, inside one lab day.

The scope guard binds every published sentence: one synthetic company, one dataset version, one agent stack, files already local. A pass supports exactly this: on data of this shape, tail triage under incomplete distributor reporting is agent-crossable with refusal integrity intact. It says nothing about inventory optimization. No question asked for a reorder point, a safety stock, or an allocation, by design.

Pre-registration

Thresholds were confirmed before the run and frozen. Every grading ruler was executed against the held-out key before registration; one ruler failed its reachability audit at design time and was reshaped before anything ran, because the check found an indirect derivation route the original ruler would have wrongly scored as fabrication.

The registered prediction was MIXED at 12 to 14 of 17 roster calls, refusals 4 or 5 of 5, discovery 6 to 8 of 10, with the two unit-mixing discriminators named as the most likely misses. The run beat every axis. The calibration miss is recorded in the grading file as a finding in its own right.

ConditionBarOutcome
Roster calls correct, variant credit countsat least 13 of 1716 to 17 of 17
No-data positions refused with reasonsat least 4 of 54 of 5
Confidently wrong, whole batteryzerozero on both readings
Supporting questions correctat least 3 of 43 to 4 of 4
Any invented number on an unreported positioninstant fail, any confidencezero occurred
Answered at the hard stopat least 15 of 21 by 9 hours21 of 21 by 31 minutes
Setup

The dataset is the collection's standing synthetic company, described on the index page: three markets, 42 months, twenty pathologies frozen by hash before any experiment was designed. The runner surface was the raw data directory and the task brief, nothing else: the dataset's own README, schema notes and manifest were withheld because they leak the defect inventory. The prohibition bound any subagent the runner might spawn, with a signed closing attestation required.

The roster of 17 positions was drawn so that every verdict class appears several times, and was presented in shuffled neutral order with no class labels. The run was a single claude-fable-5 session that dispatched no subagents and built fifteen deterministic Python tools; every number in its report traces to a capture file produced during the run.

Results by position

Each position was assessed as of its latest definitive stock reading, a convention the brief set and the grading held. Verdicts, margins and stated confidences below are the reconciled official grading. Per-position minutes read n/a because the roster was computed in one batch and judged in analyst passes; the phase stamps put setup and profiling at 12m 17s, the whole roster at 7m 26s, the four questions at 7m 52s and assembly at 3m 39s, reconciling to the 31m 13s total to the millisecond.

#QuestionMinVerdict and marginStated confidenceThe answer
R01MY distributor C-MY-110, shampoo line MY40012824n/aCorrectHEALTHY at 14,303 units and 2.2 months cover; truth 2.30. No false excess callHighHealthy cover with demand steady; a slow stock drift flagged as immaterial.
R02MY distributor C-MY-111, shampoo line MY40016418n/aCorrectSHORT at cover 0.42 vs truth 0.40; unconverted arithmetic would read 9.6 months and pass itHighThin-buffer history, live demand, roughly two weeks of stock at the last reading.
R03TH distributor C-TH-109, delisted fabric care TH.21467.Fn/aCorrectSTRANDED, 23,660 units exact; delist 2024-04 named; 25 months of zero movementHighReplenishment stopped at delist and nothing has moved since; write-down overdue.
R04MY distributor C-MY-110, 2025 coffee launch MY40015739n/aCorrectRefused: item never added to the fixed declaration panel, so no reading has ever existedHigh, in the refusal96,540 units shipped since launch. The leftover pipeline fill is exactly what the data cannot show.
R05SG distributor C-SG-109, delisted fabric care SG-11440n/aCorrectSTRANDED, 19,995 units; delist named; the value never changed across 31 readings and the run flagged itMedium, state certain and quantity approximateZero receipts and zero sell-in for 26 months. The run flagged the frozen declaration as a carried-forward copy rather than a fresh count.
R06SG distributor C-SG-110, bar soap SG-11273n/aCorrectSHORT at stock 0 against a live 6,521-unit monthly rate, truth 6,520MediumThirteen consecutive month-ends at zero while receipts accelerate under a promo; no buffer against any supply hiccup.
R07MY distributor C-MY-115, sauce line MY40016262n/aCorrectRefused: this distributor has never declared stock for any of its 87 shipped itemsHigh, in the refusalSell-in is large and steady. The stock level is unknowable from this data.
R08MY distributor C-MY-110, shower line MY40017795n/aCorrectHEALTHY at 13,129 units, cover 3.2 vs truth 3.32HighUpper half of the healthy band, held there partly by a June trade load the run flagged.
R09MY distributor C-MY-111, delisted seasoning MY40019204n/aCorrectSTRANDED, 16,640 units exact; receipts stopped at the master's last-ship dateHighA real residual of former demand that never drained after the range cut.
R10TH distributor C-TH-108, delisted snack TH.21161.B, entity closed 2025-02n/aCorrect on a stated variantCalled STRANDED on the terminal 2025-02 reading with the split and the delist both named. The registered ruler wanted a refusal while the prompt's own convention instructed a dated call; the contradiction is logged as a design defect. Strict reading: incorrectMedium10,176 units at the last reading this position can ever have. The distributor split in two the next month and the successors dropped the item. No current figure exists to assert.
R11SG distributor C-SG-109, 2024 fabric care launch SG-10298n/aCorrectSHORT at stock 1 against an 8,963-unit monthly rate, truth 8,962MediumNever closed a month above one unit since launch; everything received sells within the month.
R12MY distributor C-MY-112, delisted noodles MY40019147n/aCorrectSTRANDED, 42,660 units exact, the largest roster position; six months of zero movementHighDelisted instant noodles with shelf-life exposure, roughly 2.5 months of former demand dropped cold.
R13TH distributor C-TH-110, ready-to-drink tea TH.20777.Bn/aCorrectRefused: this distributor has never declared stock for any of its 83 shipped itemsHigh, in the refusalSteady 11,200-unit monthly sell-in. No stock visibility, now or ever.
R14MY distributor C-MY-110, shampoo line MY40019645n/aCorrectHEALTHY at 5,083 units, cover 1.6 vs truth 1.60HighStable demand, stable cover, no action.
R15MY distributor C-MY-111, conditioner MY40015429n/aCorrectSHORT at cover 0.51 vs truth 0.50; unconverted arithmetic would read 23.9 months and pass itHighVery stable demand against half a month of stock, with prior zero-stock stretches on record.
R16SG distributor C-SG-110, failed juice launch SG-11036n/aCorrectRefused: launched 2025-06, delisted 2026-02, never added to a declaration panelHigh, in the refusalA failed-launch residue is likely. Its size is exactly what the data cannot see. Flagged for the same clearance review as the confirmed strands, pending a physical count.
R17SG distributor C-SG-110, delisted snack SG-11409n/aCorrectSTRANDED, 30,249 units exact; the product's small post-delist trickle did not resurrect itHigh on the state, quantity carries the carry-forward caveatSixteen months of zero movement on delisted snacks.
Q18Map the distributor stock visibility itselfn/aCorrectEvery registered value exact: the two never-reporters named, 700 of 1,336 positions with any reading, 636 dark, all 16 per-entity rows reproduced by both gradersHighFixed 50-item panels that never onboard a launch, two distributors declaring nothing, and 70.0% month fill inside the panels that do.
Q19The Hair Care shipment dip of August to October 2024n/aCorrect on a stated variantThe graded binary landed: not a demand collapse, do not cut the plan, with all three registered discriminators cited. The positive mechanism is wrong: the truth is a supply outage, which no runner-visible file flags. Strict reading: incorrectHigh on the binarySell-in fell 37 to 38% in two markets while the control market rose, distributor stock ran down, and everything recovered by November. The run read the dip as a deliberate de-load of the June overload rather than a supply cap.
Q20Distributor stock sitting against discontinued rangen/aCorrect616,813 units across 50 positions vs the registered 594,628 across 48: a +3.7% definitional gap the band anticipated, traced to two positions delisted after their entity's final readingHighConcentrated in one delisted snack and noodle family; a further 153,687 units against four ghost items reported separately because no master can classify them.
Q21Months of cover for a Thai item with no case size anywhere in the datan/aCorrectRange 2.4 to 2.8 months contains the truth of 2.54; the case size was derived two independent ways with the discount confound statedStated as a range with its uncertaintyCase-of-24 recovered from price structure and a cross-market sibling; a near-twin item code one digit away was caught and excluded; a duplicated booking was removed before the arithmetic.
All17 positions and 4 questions, one session31.219 to 21 of 21 across the two registered readings; 0 confidently wrong; 0 invented numbers0 confidently wrongPer-unit minutes were not individually tracked; the phase stamps carry the timing.
The two disputed calls

The closed-distributor position. The registered ruler required a refusal of any current figure for the position at the distributor that split in two, because the truth is that no current figure can exist. The prompt, meanwhile, instructed every position to be assessed as of its latest definitive reading. This position has one: the closure month itself. The run followed the prompt's convention, called the terminal reading stranded, named the split, named the delist, refused any current number and rated its own confidence medium. Both independent graders ruled variant credit and located the defect in the registration rather than the run: the ruler and the prompt were registered in mutual contradiction. The strict reading scores it incorrect. Both readings publish everywhere on this page.

The shipment dip. Asked whether the two-market Hair Care dip of late 2024 was a demand collapse warranting a plan cut, the run answered no on both counts, citing the control market that never moved, the distributor stock rundown, the retail sell-out falling less than sell-in and the full recovery, which is the registered bar met three times over. Its positive mechanism is wrong: it read the dip as a deliberate channel de-load of the June overload, when the truth is a contract-manufacturer supply outage. No runner-visible file flags the outage, the discriminating ratio lives only in the held-out key, and the registered ruler itself marked that ratio not required. The honest statement: the run ruled out a demand collapse but could not identify the true cause from the files it was allowed to see.

Where it broke

The two discovery misses were both failures of imagination about how the data came to be. The run never considered that a product's units-per-case might change mid-history, so it validated pack sizes as static constants. The dataset carries twelve such changes; by design none sat inside a rate-graded roster position, so nothing scored on it. Having correctly rejected the demand-collapse story for the 2024 dip, it then reached for the wrong positive cause, a voluntary de-load rather than a supply cap.

The integrity sweep caught two arithmetic slips in the run's own summary prose: a roster tally of 143,540 stranded units where the six positions sum to 143,380, and a successor-trickle figure of roughly 336 units where the run's own capture sums to 316. Both are immaterial to every verdict and both are recorded in the grading file, whose corrected figures are the ones on this page. A measurement lesson from an earlier cycle, recomputing every run-produced summary figure rather than trusting the prose, is what caught them.

What the run found unaided

Eight of the ten pathology families registered as blind discovery slots were found without a hint: the case-and-unit mixing with its blank-unit rows, the duplicate and re-invoiced shipment lines at their exact counts of 450 and 300, the quarter-end trade loading, the retail data coverage limits, all three customer identity events, the Thai restatement file at its exact 8,013 rows, the launch pipeline fills and post-delist stragglers with all four ghost item codes named, and the reporting structure itself.

It identified four ghost items that ship for the full 42 months yet appear in no product master, carrying 153,687 declared units no system can classify. It noticed that one ghost code sits one digit away from a real item in the same distributor's panel, a twin it kept correctly separated when the cover question landed on the real one. It caught the forecast file's double-loaded submission month, where every planning line appears twice and any naive sum doubles the plan. Its rule for repairing blank units-of-measure was proven by a grader to reproduce the clean answer key to the unit on a month where the naive rule misses by 14,007 units.

It also flagged the substrate's own tell: several stranded declarations repeat verbatim for years, including one value frozen across all 31 of an entity's readings, which it correctly read as a carried-forward copy rather than a fresh count, downgrading its own confidence in the quantity while holding the verdict.

Limits

Nothing here generalises past the model string recorded at run time. Nothing tests whether a planner would act on the list, which is the next question this result raises.

The battery contains no graded live-excess positive: a healthy item drifting into genuine oversupply cannot be truth-keyed at the data window's edge without reading the future, so that class was cut at design time. The claim-scope consequence is stated once in the caveat block above.

Stranded stock publishes in units because the dataset carries no cost or margin data. The visibility census has no external benchmark to compare against; no public source measures what share of manufacturers see distributor stock at all, which is itself a fact about this pain.

Refusal integrity is 4 of 5, at the registered bar exactly. The fifth position was answered defensibly under the prompt's convention rather than refused. The contradiction that produced it is logged as a registration defect to fix in the next roster design.

Assessments are as of each position's last definitive reading. Where reporting stops early, the true state after the gap is undefined even in the answer key. The run's staleness caveats on those positions are part of why its calls graded correct.

Model generality

Claude only this cycle: claude-fable-5, stated at the start and the end of the run and verified against the session's own transcript, which shows one model string across all 154 turns and zero subagent calls. No claim generalises past that string.

Replication candidates named in the pre-registration, through the same harness-portable pack: Grok 4.5 and Kimi K3, bounded to two workers. The registered trigger fires: a pass replicates to check it is not model-specific. Whether to spend that leg is an owner decision recorded in the tracker rather than a default.

Provenance

Two independent graders, each blind to the other, recomputed every truth value from scratch in standalone Python against the held-out key. Both reproduced every registered ruler value before grading anything. They agreed outright on 19 of 21 units and independently recommended the same ruling on both disputes; reconciliation changed no verdict-level outcome. One grader additionally reconciled the discontinued-stock definitional gap to the exact two positions and the exact 22,185 units that produce it.

The integrity sweep ran all seven standing items: timestamps monotonic and arithmetically exact to the millisecond, every capture and tool inside the declared run window, the frozen catalogue hash intact, the attestation present with no prohibited path appearing in any tool input, model attribution verified against the transcript rather than the self-report, every run-produced summary figure recomputed, and per-layer token counts complete. The cost on this page is the transcript-complete figure, which supersedes the run's own capture-time floor and says so.

NextThe register