Home

The Lab

Organisational Intelligence

Humans vs Agentic AI, when to use each?

Here's proof of their commercial value, using mid-sized APAC FMCGs as the base example.

PASS0110 questions on commercial performance by the managementAgentic AI answered all ten in 137 minutes at 80% accuracy and 0% confidently wrong.General Management137 min, 8 of 10 correct, $49.53 in tokensNo human comparison publishedRan 27 to 28 July 2026Read the runPASS0210 what-if questions on the commercial plan by the managementAgentic AI answered ten what-ifs in 128 minutes at 80% accuracy. It failed 1 of 2 refusal tests and missed 3 to 8 of 31 planted defects.Commercial Finance128 min, 8 of 10 correct, $60.92 in tokensNo human comparison publishedRan 29 July 2026Read the runPASS0310 questions on making the monthly numbers agreeAgentic AI reconciled six unseen months on the first attempt with no code changes, for $3.43 of tokens. 10 of 10 checkpoints passed, 0 true breaks missed, 47 items still need a person.Finance Operationsfirst attempt, 10 of 10 checkpoints, $3.43 in tokensNo human comparison publishedRan 29 to 30 July 2026Read the runFAIL0412 facts a supplier asserts across the negotiation tableAgentic AI ran 67% to 92% accuracy across 3 models. 3 of 3 asserted a wrong number at full confidence.Key Account Salesup to 11 of 12 correct, 3 confidently wrong, $56.13 in tokensNo human comparison publishedRan 30 July 2026Read the runPASS057 questions on whether the sales history reflects real demandAgentic AI found 12 of 12 contaminated months in both runs and cut forecast error by about a third. Accuracy was 83% on one run and 50% on an identical rerun.Demand Planning71 min, two runs scored 5 of 6 and 3 of 6, $83.06 in tokensNo human comparison publishedRan 30 to 31 July 2026Read the runMIXED0610 what-if questions on the commercial plan replayed on four other enginesFour engines answered the same blind pack at 50% to 80% accuracy. 4 of 4 invented the number they should have refused.Commercial Finance14 to 37 min per engine, 4 of 4 fabricated, $9.20 in tokensNo human comparison publishedRan 30 to 31 July 2026Read the runPASS0817 stock positions to triage and 4 questions on the distributor tailAgentic AI triaged all 21 units in 31 minutes at 19 to 21 of 21 correct and 0 confidently wrong, and invented no number for any of the 5 positions the data cannot see.Demand Planning31 min, 19 to 21 of 21 correct, $67.41 in tokensNo human comparison publishedRan 3 August 2026Read the runPASS097 questions on why reported net sales movedAgentic AI answered all 7 in 38 minutes at 21 to 26 of 29 checkpoints correct and 0 confidently wrong, closed every bridge, and refused all 3 planted misattribution baits.Commercial Finance38 min, 21 to 26 of 29 correct, $65.12 in tokensNo human comparison publishedRan 4 August 2026Read the run

The test company

Velari Consumer Asia is invented, and its data is generated from a fixed seed so any run can be reproduced. It is built to the shape of a mid-sized APAC consumer goods business, mess included.

Business
Manufacturer and marketer of everyday consumer goods, headquartered in Singapore
Markets
Singapore, Malaysia, Thailand
Sells
7 brands across personal wash, hair care, home care, packaged foods and beverages
Range
180 distinct products, listed 400 times across the three market systems
Customers
99 accounts across modern trade, e-commerce, distributors, general trade and food service
Route to market
Direct to retail chains and through distributors, whose stock reporting is partial
Scale
Roughly US$180M to US$190M net sales a year
Books
Trades in SGD, MYR and THB, reports in USD
Period
42 months, January 2023 to June 2026
Systems
One regional ERP, a separate finance consolidation, a second transaction system in Thailand, retailer POS feeds and planning spreadsheets
Volume
117,174 shipment lines, plus 8,013 restatement lines for Thailand

Why the data is hard

  • Each country tracks the same physical products under its own codes and category structures, with no shared key anywhere in the data.
  • The cross-reference that bridges the three countries is hand-maintained, and it carries gaps, wrong matches and stale entries.
  • Duplicate and re-issued invoice lines sit inside the sales history, inflating any naive total with nothing flagging it.
  • A block of Thai history was restated as a separate file of adjustments rather than a correction, so it is easy to double count or to miss.
  • Customer identities move over time through acquisitions, rebrands and a distributor splitting in two, and the records that link old to new are only partly filled in.
  • Blanks are written several different ways in the same column, names carry stray spacing and capitalisation, and dates are formatted differently from file to file.

20 pathologies were written down and hash-frozen before a single row was generated, so none could be retrofitted to suit a result. The full catalogue stays sealed while experiments still run blind against this dataset.

How an experiment runs

Every experiment starts from a validated pain and a pass or fail bar written down before anything runs. The dataset's shape and its known defects are described on this page. The files themselves stay private while later experiments still run blind against them, and each run is graded against an answer key the runner never sees.

  1. The answer key is held out from the runner, along with the build log and the validation logs, because both carry answer-key-adjacent detail.
  2. The messiness catalogue is frozen and its hash recorded. Any change of any kind creates a new dataset version.
  3. Grading is done by a separate session that never sees the run's reasoning, working only from the written results and the terminal captures.
  4. Self-report is a draft. The run's own numbers are re-run against the data and audited against the captures before anything is graded.
  5. The result publishes whichever way it lands. A failure is a better page than a pass in one respect: it names exactly what broke.

No human on the other side of these numbers

No experiment here publishes a speed or cost multiple against a human. An honest one needs a timed person on the other side of it. None of these has had one, so every page reports what the run did, how much of it was right and what the tokens cost, then leaves the reader to price their own alternative.

Each row still declares how a human denominator would have been built if one were used. Every row currently reads NONE. Three of them read otherwise until 1 August 2026, and what came down and why is written on each page rather than deleted.

Since cycle 07 the protocol has allowed a timed person, measured before the agent runs and before anyone has seen the answer key. No cycle has funded one yet, so every row still reads NONE. Cycle 07 itself has no row: its pain failed the testability gate before anything was registered, because no honest answer key could be built for it from the frozen dataset, so no run exists and nothing is published from it.