10 questions on commercial performance by the management
Ten questions answered in 137 minutes for $49.53 of tokens. 80% accuracy. 0% confidently wrong.
6 correct outright, 2 correct on an alternative reading the run had already published, 2 outside tolerance. Minutes are wall clock.
0 of 10 wrong answers were asserted with confidence
That is the failure mode that costs a business money. On all ten questions the agent stated the definitional fork it had hit and which side of it the answer sat on, including on the two it got wrong.
Naming the fork did real work in the score. Six answers are correct outright. Two more count because the agent had already published the alternative reading, and that was the one inside tolerance.
Stated tentatively, because one run cannot establish it: an agent that names its definitional choices turns wrong-number risk into a visible human decision. The next cycle is designed to attack that claim.
What these numbers are and are not
One synthetic company, one run, one model. The runner held a written catalogue of where the data was broken, so this tested execution against mess already named for it. Both misses rest on defects the graders found in the instrument, and both were scored against the run anyway.
137.4 minutes of wall clock for ten questions, median 10.3 each, measured inside the run. No published benchmark covers how long a business takes to answer this particular battery, and no person was ever timed against it, so this page publishes no speed multiple. The 137 minutes is an observation, compared against nothing.
The run cost at least $49.53 in API tokens at published rates. The figure is a floor: eleven dispatched subagents did much of the per-question work and their usage records were not retained. No human cost is placed against it, because no human ran this battery.
The audit layer
Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.
The pain
A general manager's read of commercial performance lags the reporting cycle by weeks, so calls get made on instinct or on stale numbers. The target is that delivery lag, where the data already exists and the answer still arrives late. The accounting close is out of scope.
Every magnitude below is a generic corporate proxy, none of them specific to Asia Pacific FMCG.
| Baseline | Value | Source |
|---|---|---|
| Median month-end close | 6.4 calendar days, bottom quartile 10 or more | APQC benchmarking, n=2,300+ organisations |
| Average data-request turnaround | 1 to 4 weeks | Oracle Decision Dilemma, n=14,000 across 17 countries |
| Decide without consulting data because access was too hard | 76% | Oracle, same study |
| Finance teams that can run a scenario in under a day | 16 to 22% | FP&A Trends surveys, 2022 to 2025 |
| Average budget cycle, flat for three years | 8.7 weeks | AFP 2026 Benchmarking Survey, n=332 |
| FP&A time spent collecting and validating data | 46% | Singapore FP&A Board, July 2025 |
What was tested
The claim under test: from the published dataset, its schema documentation and its frozen messiness catalogue alone, with no reconciliation help and no human analyst, one fresh agent session answers all ten questions inside an eight-hour day, inside pre-registered tolerances against a held-out answer key, stating the definitional choices behind its numbers.
It did not test real enterprise data access, permissioned source systems, multi-stakeholder sign-off, data whose schema the agent has never seen, or any model beyond the one recorded. A pass supports one statement: an agent can cross the analysis and reconciliation part of decision latency on data of this shape.
The runner was allowed to read the messiness catalogue, which names all twenty pathologies and which a real analyst never gets, so this measured execution against mess already handed over in writing. Cutting the other way, each question carries a non-scoring flag for whether the run demonstrated a pathology, named it from the catalogue, or missed it, and all ten came back as demonstrated.
The dataset was built awkward on purpose: three incompatible SKU code systems, three product taxonomies, fiscal calendars mismatched against calendar months, an unlabelled foreign-exchange method behind the published USD, a restated year, duplicated and re-invoiced documents, a cross-reference file with wrong and missing links, and no agreed definition of net sales.
The bar
All four pass conditions and all three fail triggers were confirmed on 27 July 2026, before anything ran. No threshold moved afterwards. Anything between them publishes as MIXED and goes up exactly as a pass would.
The prediction registered before the run was MIXED at 6 or 7 correct, with cross-market reconciliation named as the expected breaking point. The run came in at 8, reconciliation was crossed cleanly, and the one genuine failure broke on promotional scope grain. The question called the most likely outright failure came back correct.
| Condition | Threshold | Actual |
|---|---|---|
| Pass: graded correct | 7 or more of 10 | 8 |
| Pass: confidently wrong on Q08 to Q10 | 0 | 0 |
| Pass: assumptions stated | 9 or more of 10 | 10 |
| Pass: wall clock, all ten answered | 8 hours or less | 137.4 min |
| Pass: median per question | 45 min or less | 10.3 min |
| Fail trigger: graded correct | 4 or fewer | 8 |
| Fail trigger: confidently wrong on Q08 to Q10 | any | none |
| Fail trigger: answered at the 9-hour stop | fewer than 8 | 10, at 2h 17m |
Results
Elapsed time is wall clock. A further 13.5 minutes went on shared tooling before the first question, counted in the 137-minute total and attributed to no single question. Confidence is what the run logged at the time, before grading.
Two rows read "correct on a stated variant": the primary number fell outside tolerance and an alternative the run had already published fell inside it. The pre-registration credits that, flags it separately, and counts it toward the eight.
| # | Question | Min | Verdict and margin | Stated confidence | The answer |
|---|---|---|---|---|---|
| Q01 | Singapore net invoiced sales by month, 2025 | 3.9 | CorrectAnnual figure off by 0.0008%; 12 of 12 months inside 2% | High | SGD 52.35m commercial net, SGD 49.48m on the finance definition. Weakest month November. |
| Q02 | Malaysia top ten customers, 2024 against 2023 | 5.4 | CorrectTop-ten overlap 10 of 10; top-three values inside 0.004% | High | Led by Harmoni Mart at MYR 34.6m, whose +45.5% is a merger transfer. The combined book is down 7.6%. |
| Q03 | Singapore launches: pipe-fill and verdicts | 8.0 | Correct on a stated variantLaunches 8 of 8, failures 2 of 2; run-rate 5 of 8 on the primary mean, 6 of 8 on the run's stated median, bar 6 | High | Eight launches in window, two of them failed. Pipe-fill of 51k to 169k units each. |
| Q04 | Group net sales and EBIT in USD, first half 2026 | 7.9 | Outside toleranceSingapore EBIT 4.08% off against a 3% band, the breach that scored this wrong; Malaysia EBIT 3.06% primary, 1.67% on the run's variant | High on totals, medium on Malaysia | Net USD 85.9m, EBIT USD 9.8m at monthly-average FX. Malaysia's June has a three-day hole in the P&L, reported as one. |
| Q05 | Malaysia-only lines and true three-market products | 13.5 | CorrectMalaysia-only overlap 9 of 10, bar 7; three-market set 5 of 5, bar 4 | Medium | Ten Malaysia-only drivers worth 4.61m units, half of them pack-size gaps. Biggest three-market line is Rivo Lemon Tea 1L at 1.81m units. |
| Q06 | The nine largest promotions of 2025 | 16.4 | Outside tolerance2 of 9 inside the 15% band at the pre-registered market grain, bar 6; payback direction right 9 of 9 | Medium | Incremental 46k to 182k units each, payback 12% to 43%. Between 30% and 44% of trade spend attributes to no promotion at all. |
| Q07 | Distributor loading, eight quarters | 22.9 | CorrectMalaysia 5.27%, Thailand 1.18%, both inside 15%; 6 of 6 loading months named exactly, no false positive on June 2026 | Medium-high | Loading of 3.42m units, 4.4% of distributor sell-in, confined to June and December. No material payback exposure into July 2026. |
| Q08 | Thailand after the rebate restatement | 6.7 | Correct4 of 4 quarters inside 2%; restatement magnitude exact to four decimal places; the double-count trap did not fire | High | The restatement is THB 40.18m, 2.0% of net. Thai calendar 2024 now reads THB 1,996.6m commercial and 1,912.1m finance. |
| Q09 | Group net revenue, twelve months to June 2026 | 12.6 | Correct on a stated variantBoth definitions inside 2%; the bridge missed at 5.94% primary and holds at 0.20% on the run's variant | High | Use the finance number, USD 172.1m. Commercial is USD 182.1m, and the 5.8% gap is off-invoice trade spend. |
| Q10 | The Hair Care fall of late 2024 | 23.0 | CorrectLost volume 6.1% to 12.6% off in Malaysia and 1.2% to 2.4% in Thailand, both inside the 20% band | High on diagnosis, medium on size | A supply constraint from August to October 2024. Lost sell-in of MYR 14.7m and THB 89.7m. |
| All | Ten questions, one session | 137.4 | 8 of 10 correct (2 on a stated variant), 2 outside tolerance | 0 confidently wrong | Median 10.3 minutes per question. |
Where it broke
Question 04, Singapore EBIT. That band cannot be reached from the files the runner was allowed to see: the published management P&L diverges from the accrual-true answer key by up to 12.5% in a single month, by construction. One grader rebuilt the run's answer from runner-visible data alone and reproduced it to the dollar on all eight reported aggregates, so the run sat at the achievable ceiling. Scored wrong anyway, and the band is flagged for the next dataset version.
Question 06, promotional uplift. The answer key holds volume at market, product and month, a grain that cannot isolate a promotion run with a single customer: aggregating over a promotion's scope sweeps in other customers' volume, at between 0% and 118% contamination. Recomputed at the customer grain the promotions actually ran at, 8 of 9 land inside tolerance. Both readings are on the record.
Neither defect moved the score, because the official count follows the rule locked before the run. Grading a run more kindly once you have seen its answers is how a lab stops being a lab.
The diagnostic worth keeping
The dataset hands over a deliberately broken cross-market product cross-reference: links missing, links pointing at the wrong counterpart, stale codes, and a free-text confidence signal. It is the only bridge between the three markets, and the pre-registration named the question that depends on it as the most likely outright failure.
Both graders agree on what the run did with it. The identity links it validated graded 100% precise, it caught all 21 truly wrong links, and it discarded zero good ones. The question those links fed came back correct.
One figure is deliberately absent. The graders' face-value precision readings for the file itself disagree on how proposed links should be counted while agreeing on the underlying truth, so neither the file's link total nor either side's precision reading appears here.
The run logged only medium confidence on its own answer, and it was right to: an incomplete cross-reference cannot prove a product has no counterpart anywhere.
Limits
No repeat run, so no variance estimate.
Two of the ten scores rest on instruments the graders themselves flagged as defective. Both were scored against the run, the conservative direction, so 8 of 10 is the strict-letter reading. A corrected instrument would move the score up, which is a reason to distrust the test as much as the result.
No human ran this battery, so this page states nothing about how much faster the agent was. Until 1 August 2026 it published a 28-to-52 times multiple against a bottom-up estimate of 8 to 15 analyst days. That estimate had no public benchmark under it and no timed person behind it. The pre-registration marked it an assumption from the start, so the multiple was withdrawn rather than defended.
The agent also started with every file already on disk, so it skipped the access queue that carries most of the delay a business actually feels. Any comparison built later has to charge the agent for that queue or exclude it from the human side.
Correct against an answer key is not the same as trusted by a general manager. Nothing here tests whether a P&L owner would sign their name under these numbers, which is the brief for the next cycle.
Model generality
Claude only this cycle. The exact model string is printed at the top of this page, and no claim generalises past it.
Replication candidates named in the pre-registration, through the same pack and the same questions: Grok 4.5, Kimi K3, Qwen3.7 Max. A pass gets replicated to check it is not an artefact of one model family.
Provenance
Thresholds pre-registered and confirmed before the run. The answer key, the build log and the validation logs were all held out from the runner throughout.
Two independent graders scored the run, each blind to the other's work, neither running the run's own tooling. Each recomputed every truth value from scratch against the held-out key, so no number here was taken on the run's word. Where they disagreed, the disagreement is recorded and the reconciliation ruling is published alongside it.
Timings were independently audited. All 30 capture files and all 13 tool files fall strictly inside the question windows the run declared, ordering is strictly monotonic, no discrepancy exceeds two minutes, and neither grader found evidence of retrospective stamping. Captures include the failed and superseded runs, and nothing was reconstructed afterwards.
One procedural gap is on the record. The run omitted the closing attestation its instructions required. Its substance was verified independently, no held-out path touched and no out-of-scope file modified in the run window, and the gap does not affect the score.