Organisational Intelligence
Experiment 01PASSRan 27 to 28 July 2026Model claude-fable-5

10 questions on commercial performance by the management

Ten questions answered in 137 minutes for $49.53 of tokens. 80% accuracy. 0% confidently wrong.

137 minTen questions, one session
8 / 10Inside tolerance, with no wrong answer asserted confidently
$49.53API tokens, the run only
10.3 minMedian per question
0Confidently wrong answers
10 / 10Assumptions stated

6 correct outright, 2 correct on an alternative reading the run had already published, 2 outside tolerance. Minutes are wall clock.

0 of 10 wrong answers were asserted with confidence

That is the failure mode that costs a business money. On all ten questions the agent stated the definitional fork it had hit and which side of it the answer sat on, including on the two it got wrong.

Naming the fork did real work in the score. Six answers are correct outright. Two more count because the agent had already published the alternative reading, and that was the one inside tolerance.

Stated tentatively, because one run cannot establish it: an agent that names its definitional choices turns wrong-number risk into a visible human decision. The next cycle is designed to attack that claim.

What these numbers are and are not

One synthetic company, one run, one model. The runner held a written catalogue of where the data was broken, so this tested execution against mess already named for it. Both misses rest on defects the graders found in the instrument, and both were scored against the run anyway.

137.4 minutes of wall clock for ten questions, median 10.3 each, measured inside the run. No published benchmark covers how long a business takes to answer this particular battery, and no person was ever timed against it, so this page publishes no speed multiple. The 137 minutes is an observation, compared against nothing.

The run cost at least $49.53 in API tokens at published rates. The figure is a floor: eleven dispatched subagents did much of the per-question work and their usage records were not retained. No human cost is placed against it, because no human ran this battery.

The audit layer

Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.

The pain

A general manager's read of commercial performance lags the reporting cycle by weeks, so calls get made on instinct or on stale numbers. The target is that delivery lag, where the data already exists and the answer still arrives late. The accounting close is out of scope.

Every magnitude below is a generic corporate proxy, none of them specific to Asia Pacific FMCG.

BaselineValueSource
Median month-end close6.4 calendar days, bottom quartile 10 or moreAPQC benchmarking, n=2,300+ organisations
Average data-request turnaround1 to 4 weeksOracle Decision Dilemma, n=14,000 across 17 countries
Decide without consulting data because access was too hard76%Oracle, same study
Finance teams that can run a scenario in under a day16 to 22%FP&A Trends surveys, 2022 to 2025
Average budget cycle, flat for three years8.7 weeksAFP 2026 Benchmarking Survey, n=332
FP&A time spent collecting and validating data46%Singapore FP&A Board, July 2025
What was tested

The claim under test: from the published dataset, its schema documentation and its frozen messiness catalogue alone, with no reconciliation help and no human analyst, one fresh agent session answers all ten questions inside an eight-hour day, inside pre-registered tolerances against a held-out answer key, stating the definitional choices behind its numbers.

It did not test real enterprise data access, permissioned source systems, multi-stakeholder sign-off, data whose schema the agent has never seen, or any model beyond the one recorded. A pass supports one statement: an agent can cross the analysis and reconciliation part of decision latency on data of this shape.

The runner was allowed to read the messiness catalogue, which names all twenty pathologies and which a real analyst never gets, so this measured execution against mess already handed over in writing. Cutting the other way, each question carries a non-scoring flag for whether the run demonstrated a pathology, named it from the catalogue, or missed it, and all ten came back as demonstrated.

The dataset was built awkward on purpose: three incompatible SKU code systems, three product taxonomies, fiscal calendars mismatched against calendar months, an unlabelled foreign-exchange method behind the published USD, a restated year, duplicated and re-invoiced documents, a cross-reference file with wrong and missing links, and no agreed definition of net sales.

The bar

All four pass conditions and all three fail triggers were confirmed on 27 July 2026, before anything ran. No threshold moved afterwards. Anything between them publishes as MIXED and goes up exactly as a pass would.

The prediction registered before the run was MIXED at 6 or 7 correct, with cross-market reconciliation named as the expected breaking point. The run came in at 8, reconciliation was crossed cleanly, and the one genuine failure broke on promotional scope grain. The question called the most likely outright failure came back correct.

ConditionThresholdActual
Pass: graded correct7 or more of 108
Pass: confidently wrong on Q08 to Q1000
Pass: assumptions stated9 or more of 1010
Pass: wall clock, all ten answered8 hours or less137.4 min
Pass: median per question45 min or less10.3 min
Fail trigger: graded correct4 or fewer8
Fail trigger: confidently wrong on Q08 to Q10anynone
Fail trigger: answered at the 9-hour stopfewer than 810, at 2h 17m

Results

Elapsed time is wall clock. A further 13.5 minutes went on shared tooling before the first question, counted in the 137-minute total and attributed to no single question. Confidence is what the run logged at the time, before grading.

Two rows read "correct on a stated variant": the primary number fell outside tolerance and an alternative the run had already published fell inside it. The pre-registration credits that, flags it separately, and counts it toward the eight.

#QuestionMinVerdict and marginStated confidenceThe answer
Q01Singapore net invoiced sales by month, 20253.9CorrectAnnual figure off by 0.0008%; 12 of 12 months inside 2%HighSGD 52.35m commercial net, SGD 49.48m on the finance definition. Weakest month November.
Q02Malaysia top ten customers, 2024 against 20235.4CorrectTop-ten overlap 10 of 10; top-three values inside 0.004%HighLed by Harmoni Mart at MYR 34.6m, whose +45.5% is a merger transfer. The combined book is down 7.6%.
Q03Singapore launches: pipe-fill and verdicts8.0Correct on a stated variantLaunches 8 of 8, failures 2 of 2; run-rate 5 of 8 on the primary mean, 6 of 8 on the run's stated median, bar 6HighEight launches in window, two of them failed. Pipe-fill of 51k to 169k units each.
Q04Group net sales and EBIT in USD, first half 20267.9Outside toleranceSingapore EBIT 4.08% off against a 3% band, the breach that scored this wrong; Malaysia EBIT 3.06% primary, 1.67% on the run's variantHigh on totals, medium on MalaysiaNet USD 85.9m, EBIT USD 9.8m at monthly-average FX. Malaysia's June has a three-day hole in the P&L, reported as one.
Q05Malaysia-only lines and true three-market products13.5CorrectMalaysia-only overlap 9 of 10, bar 7; three-market set 5 of 5, bar 4MediumTen Malaysia-only drivers worth 4.61m units, half of them pack-size gaps. Biggest three-market line is Rivo Lemon Tea 1L at 1.81m units.
Q06The nine largest promotions of 202516.4Outside tolerance2 of 9 inside the 15% band at the pre-registered market grain, bar 6; payback direction right 9 of 9MediumIncremental 46k to 182k units each, payback 12% to 43%. Between 30% and 44% of trade spend attributes to no promotion at all.
Q07Distributor loading, eight quarters22.9CorrectMalaysia 5.27%, Thailand 1.18%, both inside 15%; 6 of 6 loading months named exactly, no false positive on June 2026Medium-highLoading of 3.42m units, 4.4% of distributor sell-in, confined to June and December. No material payback exposure into July 2026.
Q08Thailand after the rebate restatement6.7Correct4 of 4 quarters inside 2%; restatement magnitude exact to four decimal places; the double-count trap did not fireHighThe restatement is THB 40.18m, 2.0% of net. Thai calendar 2024 now reads THB 1,996.6m commercial and 1,912.1m finance.
Q09Group net revenue, twelve months to June 202612.6Correct on a stated variantBoth definitions inside 2%; the bridge missed at 5.94% primary and holds at 0.20% on the run's variantHighUse the finance number, USD 172.1m. Commercial is USD 182.1m, and the 5.8% gap is off-invoice trade spend.
Q10The Hair Care fall of late 202423.0CorrectLost volume 6.1% to 12.6% off in Malaysia and 1.2% to 2.4% in Thailand, both inside the 20% bandHigh on diagnosis, medium on sizeA supply constraint from August to October 2024. Lost sell-in of MYR 14.7m and THB 89.7m.
AllTen questions, one session137.48 of 10 correct (2 on a stated variant), 2 outside tolerance0 confidently wrongMedian 10.3 minutes per question.
Where it broke

Question 04, Singapore EBIT. That band cannot be reached from the files the runner was allowed to see: the published management P&L diverges from the accrual-true answer key by up to 12.5% in a single month, by construction. One grader rebuilt the run's answer from runner-visible data alone and reproduced it to the dollar on all eight reported aggregates, so the run sat at the achievable ceiling. Scored wrong anyway, and the band is flagged for the next dataset version.

Question 06, promotional uplift. The answer key holds volume at market, product and month, a grain that cannot isolate a promotion run with a single customer: aggregating over a promotion's scope sweeps in other customers' volume, at between 0% and 118% contamination. Recomputed at the customer grain the promotions actually ran at, 8 of 9 land inside tolerance. Both readings are on the record.

Neither defect moved the score, because the official count follows the rule locked before the run. Grading a run more kindly once you have seen its answers is how a lab stops being a lab.

The diagnostic worth keeping

The dataset hands over a deliberately broken cross-market product cross-reference: links missing, links pointing at the wrong counterpart, stale codes, and a free-text confidence signal. It is the only bridge between the three markets, and the pre-registration named the question that depends on it as the most likely outright failure.

Both graders agree on what the run did with it. The identity links it validated graded 100% precise, it caught all 21 truly wrong links, and it discarded zero good ones. The question those links fed came back correct.

One figure is deliberately absent. The graders' face-value precision readings for the file itself disagree on how proposed links should be counted while agreeing on the underlying truth, so neither the file's link total nor either side's precision reading appears here.

The run logged only medium confidence on its own answer, and it was right to: an incomplete cross-reference cannot prove a product has no counterpart anywhere.

Limits

No repeat run, so no variance estimate.

Two of the ten scores rest on instruments the graders themselves flagged as defective. Both were scored against the run, the conservative direction, so 8 of 10 is the strict-letter reading. A corrected instrument would move the score up, which is a reason to distrust the test as much as the result.

No human ran this battery, so this page states nothing about how much faster the agent was. Until 1 August 2026 it published a 28-to-52 times multiple against a bottom-up estimate of 8 to 15 analyst days. That estimate had no public benchmark under it and no timed person behind it. The pre-registration marked it an assumption from the start, so the multiple was withdrawn rather than defended.

The agent also started with every file already on disk, so it skipped the access queue that carries most of the delay a business actually feels. Any comparison built later has to charge the agent for that queue or exclude it from the human side.

Correct against an answer key is not the same as trusted by a general manager. Nothing here tests whether a P&L owner would sign their name under these numbers, which is the brief for the next cycle.

Model generality

Claude only this cycle. The exact model string is printed at the top of this page, and no claim generalises past it.

Replication candidates named in the pre-registration, through the same pack and the same questions: Grok 4.5, Kimi K3, Qwen3.7 Max. A pass gets replicated to check it is not an artefact of one model family.

Provenance

Thresholds pre-registered and confirmed before the run. The answer key, the build log and the validation logs were all held out from the runner throughout.

Two independent graders scored the run, each blind to the other's work, neither running the run's own tooling. Each recomputed every truth value from scratch against the held-out key, so no number here was taken on the run's word. Where they disagreed, the disagreement is recorded and the reconciliation ruling is published alongside it.

Timings were independently audited. All 30 capture files and all 13 tool files fall strictly inside the question windows the run declared, ordering is strictly monotonic, no discrepancy exceeds two minutes, and neither grader found evidence of retrospective stamping. Captures include the failed and superseded runs, and nothing was reconstructed afterwards.

One procedural gap is on the record. The run omitted the closing attestation its instructions required. Its substance was verified independently, no held-out path touched and no out-of-scope file modified in the run window, and the gap does not affect the score.

NextThe register