The Lab
Organisational Intelligence
Humans vs Agentic AI, when to use each?
Here's proof of their commercial value, using mid-sized APAC FMCGs as the base example.
The test company
Velari Consumer Asia is invented, and its data is generated from a fixed seed so any run can be reproduced. It is built to the shape of a mid-sized APAC consumer goods business, mess included.
- Business
- Manufacturer and marketer of everyday consumer goods, headquartered in Singapore
- Markets
- Singapore, Malaysia, Thailand
- Sells
- 7 brands across personal wash, hair care, home care, packaged foods and beverages
- Range
- 180 distinct products, listed 400 times across the three market systems
- Customers
- 99 accounts across modern trade, e-commerce, distributors, general trade and food service
- Route to market
- Direct to retail chains and through distributors, whose stock reporting is partial
- Scale
- Roughly US$180M to US$190M net sales a year
- Books
- Trades in SGD, MYR and THB, reports in USD
- Period
- 42 months, January 2023 to June 2026
- Systems
- One regional ERP, a separate finance consolidation, a second transaction system in Thailand, retailer POS feeds and planning spreadsheets
- Volume
- 117,174 shipment lines, plus 8,013 restatement lines for Thailand
Why the data is hard
- Each country tracks the same physical products under its own codes and category structures, with no shared key anywhere in the data.
- The cross-reference that bridges the three countries is hand-maintained, and it carries gaps, wrong matches and stale entries.
- Duplicate and re-issued invoice lines sit inside the sales history, inflating any naive total with nothing flagging it.
- A block of Thai history was restated as a separate file of adjustments rather than a correction, so it is easy to double count or to miss.
- Customer identities move over time through acquisitions, rebrands and a distributor splitting in two, and the records that link old to new are only partly filled in.
- Blanks are written several different ways in the same column, names carry stray spacing and capitalisation, and dates are formatted differently from file to file.
20 pathologies were written down and hash-frozen before a single row was generated, so none could be retrofitted to suit a result. The full catalogue stays sealed while experiments still run blind against this dataset.
How an experiment runs
Every experiment starts from a validated pain and a pass or fail bar written down before anything runs. The dataset's shape and its known defects are described on this page. The files themselves stay private while later experiments still run blind against them, and each run is graded against an answer key the runner never sees.
- The answer key is held out from the runner, along with the build log and the validation logs, because both carry answer-key-adjacent detail.
- The messiness catalogue is frozen and its hash recorded. Any change of any kind creates a new dataset version.
- Grading is done by a separate session that never sees the run's reasoning, working only from the written results and the terminal captures.
- Self-report is a draft. The run's own numbers are re-run against the data and audited against the captures before anything is graded.
- The result publishes whichever way it lands. A failure is a better page than a pass in one respect: it names exactly what broke.
No human on the other side of these numbers
No experiment here publishes a speed or cost multiple against a human. An honest one needs a timed person on the other side of it. None of these has had one, so every page reports what the run did, how much of it was right and what the tokens cost, then leaves the reader to price their own alternative.
Each row still declares how a human denominator would have been built if one were used. Every row currently reads NONE. Three of them read otherwise until 1 August 2026, and what came down and why is written on each page rather than deleted.
Since cycle 07 the protocol has allowed a timed person, measured before the agent runs and before anyone has seen the answer key. No cycle has funded one yet, so every row still reads NONE. Cycle 07 itself has no row: its pain failed the testability gate before anything was registered, because no honest answer key could be built for it from the frozen dataset, so no run exists and nothing is published from it.