Organisational Intelligence
Experiment 04FAILRan 30 July 2026Model claude-fable-5, replicated on grok-4.5 and kimi-k3

12 facts a supplier asserts across the negotiation table

Twelve negotiation facts built blind for $56.13 of tokens. 67% to 92% accuracy across 3 models. 3 of 3 asserted a wrong number at full confidence.

71 minPrimary run, all twelve facts answered
11 / 12Facts correct on the best runs, with one wrong fact asserted at full confidence
$56.13API tokens, the Claude run only
claude-fable-5primary run, dual-graded
11 of 12 correctThe wrong fact: the disputed sell-out, +26.6%
grok-4.5replication
8 of 12 correctThe wrong fact: a product value inflated 8.1% by an undetected duplicate
kimi-k3replication
11 of 12 correctThe wrong fact: the disputed sell-out, +26.6%, same number as Claude
Twelve pre-registered negotiation facts per run, identical blind pack, graded against the same held-out key. Correct includes facts credited on an explicitly stated variant (two for Kimi). The marked segment is the confidently wrong fact: outside tolerance, no caveat on the failing dimension, asserted at high confidence. The pre-registered bar was zero.
3 of 3Models failed the zero-confidently-wrong bar
+26.6%Size of the shared wrong fact, a sell-out the buyer's own scan file catches
12 of 14Hidden data pathologies found unaided by the best runs, blind
0 of 3Models that noticed the scan article covers two products

11 correct outright, 0 correct on an alternative reading the run had already published, 1 outside tolerance and confidently wrong. Minutes are wall clock.

All 3 models failed on the fact only the counterparty's data can check

The designed trap was Q05. The buyer disputes our sell-out claim on one product, and in their scan file that product shares a single retailer article with a sibling SKU: the scan series is both products added together. All three models matched the article by barcode, concluded the identity was proven, and asserted the merged total of 216,259 units as the disputed product's sell-out at high confidence. The true figure is 170,783. Each model had already computed the number that falsifies its own answer, a scan total running 27% above the same year's sell-in, and each explained it away.

Grok failed differently as well. Its confidently wrong fact was a top-five product value inflated 8.1% by a duplicated invoice line it never detected, in an answer whose data-quality note said nothing material was found. So the wrong fact changes from model to model, and the behaviour that produces it does not: a clean-looking number, a plausible identity story, and full confidence.

The pattern worth carrying into any negotiation prep: the facts a supplier can verify against its own ledgers came out almost entirely right, to the cent in several cases. The facts that broke are the ones only the counterparty's data can check. Zero of the three models considered that the retailer's article might cover two products, and the flag questions show the honest reflex exists elsewhere: two of three runs surfaced the two-definitions split on net revenue unprompted, and all three declined to assert a consumption figure the data cannot support.

What these numbers are and are not

One synthetic company, one run per model, graded once against a pre-registered answer key. The runs were blind: no schema documentation, no list of what was broken in the data. Three of the twelve questions were designed so the correct answer is a flag rather than a number, and one was a designed trap on a product the retailer's scan file merges with a sibling. Replication runs were graded by one independent grader each against the same key; the primary run got two.

The buyer side of the table is sourced: McKinsey finds 75% of large buyers spend 10 or more days preparing a large negotiation, and Bain finds the best retailers arrive with SKU-level cross-channel comparisons. No public figure exists for the seller side, so this experiment ships no speed claim. The primary run took 71 minutes; that number is an observation, compared against nothing.

The Claude run cost at least $56.13 in API tokens at published rates. The figure is a floor, because one dispatched subagent was never separately metered. The two replication engines were metered directly by the gateway, at $1.77 for grok-4.5 and $3.98 for kimi-k3, over runs of 25 and 44 minutes against the primary's 71. No human cost is placed against any of them.

The audit layer

Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.

Why this pain is real

The validated pain, Q4 in this lab's queue: JBP and negotiation prep is manually rebuilt each cycle and produces a weaker, less defensible account-level fact base than the retailer's buying team brings, so the plan gets picked apart rather than jointly owned.

The buyer side is well evidenced. McKinsey's key-account work finds 75% of large buyers spend 10 or more days preparing per large negotiation, and 60% centralise purchasing for better terms. Bain's key-account research finds the best retailers bring SKU-level, cross-channel comparisons to the table. The counter-evidence is carried too: trade press has documented JBPs that are supplier-authored theatre, and a mature vendor ecosystem already sells against this pain.

The seller side is a named gap: no public figure exists for hours per JBP cycle or for the error rate of manually built fact bases. That is why this experiment measures defensibility and ships no speed claim.

The sub-claim, pre-registered

Given only the raw data files, no schema documentation and no catalogue of known defects, a single fresh agent session can assemble account fact packs for three accounts spanning the dataset's identity pathologies, a merger, a rename and a distributor split, answering 12 negotiation-grade questions such that the asserted facts hold against a held-out key, unsupportable facts are flagged rather than asserted, and zero facts are confidently wrong, inside one lab day.

Scope guard, binding on every sentence here: one synthetic company, one dataset version, one run per model. This does not test live negotiation dynamics, real enterprise data access, or any model not named on this page. A pass would have supported exactly one statement: on data of this shape, the fact-base-assembly portion of JBP prep is agent-crossable at negotiation-grade defensibility.

The pre-registered bar

Pass required all four, written down and confirmed before the run: at least 9 of 12 facts correct; zero confidently wrong facts across all twelve; at least 2 of the 3 flag questions correctly flagged; assumptions stated on at least 11 of 12. Fail fired on any one of: a single confidently wrong fact anywhere, 6 or fewer correct, or fewer than 10 answered at the hard stop.

The zero-tolerance trigger is the experiment's point. One confidently wrong fact is one embarrassed negotiator, so there is no acceptable count above zero. The primary run failed on that trigger alone, with three of the four pass conditions met.

Every tolerance band traces to a grading recipe executed before registration. Three rulers failed their reachability checks during design and were reshaped before anything ran: the loading magnitude was re-based to a unit-free share because Thailand's case-to-unit conversion factor exists nowhere in the runner-visible data, the share bands went dual-path for the same reason, and the net-revenue question was re-anchored from an account where the definitional gap was 0.14% to one where it is 10.2%. The cycle-1 lesson, applied: no known-defective ruler enters the pre-registration.

The graded record

The per-question record below is the primary run. Verification order: an integrity pass reconciled every self-reported number to its terminal capture and re-ran three of the run's tools byte-identical; two independent graders then recomputed every truth value from the held-out key in separate standalone code, blind to each other; a reconciliation pass ruled on the judgment calls. Both graders reached identical verdicts on all twelve questions.

#QuestionMinVerdict and marginStated confidenceThe answer
Q01Harmoni Mart account net sales, calendar 20253.0Correct+0.053% against a 2% band; the entire gap is one undetected re-invoiced line worth MYR 26,166HighMYR 49,399,150.79 invoiced net, duplicates removed and disclosed, basis stated.
Q02Harmoni growth, 2025 versus 2024, basis the buyer's CFO would accept1.0Correct0.14pp against a 1.5pp band on the consolidated basis; the 31pp trap between bases was named and defusedHigh+11.6% with the acquired chain consolidated into the base year; the same-id +42.7% shown and rejected as indefensible.
Q03Top five products to the account, with values1.0Correct5 of 5 products, values exact to the centHighMYR 12.0m across the five, 24.3% of the account.
Q04Cross-market presence and group share for those five5.0CorrectPresence 10 of 10 flags right; shares within 0.5pp with the Thai conversion gap flagged and boundedMediumAll five presence calls right against a broken cross-reference, including one wrong mapping caught and rejected.
Q05The disputed sell-out on one named product (the designed trap)4.0Confidently wrong+26.63%: asserted the merged scan article as one product at high confidence; the split was never consideredHigh, which is the failure216,259 units asserted; the true figure for the disputed product is 170,783, and the buyer's category manager can read it off their own file.
Q06Trade investment posted to the account, calendar 20252.0CorrectExact to the cent against a 1% band, with the timing caveats carriedMediumMYR 5,072,116.50 posted, 63.7% of it unallocated provisions, posted-versus-earned divergence stated.
Q07Nakhon distributor lineage net sales, calendar 2025, restated1.0Correct+0.62% against a 2% band; restatement applied exactly once, split lineage stitchedHighTHB 212.0m across the closed entity and both successors, restated basis led with the bridge shown.
Q08Nakhon growth, 2025 versus 2024, lineage-consistent1.0Correct0.27pp against a 1.5pp bandHighEssentially flat, +1.7%, and the answer holds on either restatement basis.
Q09The quarter-end loading allegation, sized and dated2.0CorrectBoth loading months named exactly, both decoy months cleared, share inside the band by 0.24ppHighJune and December 2025, about 3.2% of the year's volume; loading shifted timing, it did not create growth.
Q10Net revenue on the account: the number we stand behind1.0CorrectBoth definitions within 0.06% of their truths, the 10.2% bridge exact to the centHighTwo honest answers carried with the bridge: invoiced net MYR 49.6m, net after trade investment MYR 44.5m.
Q11What the distributor territory actually consumed1.0CorrectCorrectly declined: no scan data exists for the territory and the inventory feed is one-third unreportedMediumRefused the precise figure, quantified the coverage gap, and answered the full-warehouses claim with their own stock trend instead.
Q12Three years of scan history across the account's rename1.0CorrectAll three years exact to the unit; a name-only filter loses 100% of two of the yearsHighOne continuous account across the rebrand, stitched by item codes and barcodes, not by name.
AllTwelve facts, one session, primary run71.011 of 12 correct, 1 outside tolerance1 confidently wrongMedian 1.0 minutes per question.
The replication

Because the primary run failed, the pre-registration required replication on non-Claude models to test whether the weakness is Claude-specific. The identical blind pack ran on Grok 4.5 and Kimi K3 in isolated copies of the repository, prior run artifacts deleted unread. It is not Claude-specific.

All three models asserted the identical wrong sell-out number, 216,259 units, at high confidence. Kimi ran the deepest identity forensics of the three, seven description spellings traced across four retailers, and still concluded the barcode proved a single product. Grok also anchored its restated-history answer on the superseded basis and named a decoy month as loading, and its confidently wrong fact came from a different trap entirely: undetected duplicate absorption.

Kimi earned two facts on explicitly stated variants, one of which came from the strongest single finding any run produced: it noticed that every unallocated trade-spend posting in the ledger is mirrored to the cent on a second customer across all 42 months, a generator artifact this lab's own design never caught. The finding was verified true and is now a logged fix for the dataset's next version.

RunVerdictCorrectConfidently wrongPathologies foundWall clock
claude-fable-5, primaryFAIL11 of 121: the disputed sell-out, +26.6%12 of 1471 min
grok-4.5, replicationFAIL8 of 121: a duplicate-inflated product value, +8.1%9 of 1425 min
kimi-k3, replicationFAIL11 of 121: the disputed sell-out, +26.6%12 of 1444 min
What the runs found unaided

The runs were blind, so finding the mess is part of what was measured. Fourteen discovery items were pre-registered, from the merger, rename and split, through duplicate and re-invoiced lines, a restated history delivered as a separate delta file, quarter-end loading with payback, trade-spend timing, and a missing unit-conversion factor. The primary run found 12 of 14, including several this lab predicted it would miss.

The two it missed are instructive. Re-invoiced companion lines, a cancelled invoice re-issued under a new document number, were tested for with the wrong signature and declared absent; the entire residual error on three money questions is exactly those lines. And the scan-article consolidation was not just missed but affirmatively ruled out, which is what armed the confidently wrong fact. No model found either one.

Deviations and grading notes

Recorded against this cycle, none scoring: the closing attestation and the runner's own unsafe-answers list were terminal-print items and are absent from the archived artifacts, so the blind condition rests on the integrity scan, zero references to any held-out file across 19 tools and 24 captures, and on behavioural evidence, since the runs walked into multiple pre-registered traps. One answer mischaracterised its own evidence in passing, calling December the year's stock peak when its cited table shows February higher. Two design-side evidence statements behind a reshaped ruler were found imprecise by a grader and corrected on the record; the reshaped rulers themselves stand.

The wall clock is reported once, here: 71 minutes primary, 25 and 44 for the replications, all twelve questions answered in each run, against an 8-hour box. No published seller-side baseline exists, so these figures are compared against nothing, and this page makes no claim from them.

Limits

One synthetic company, whose full defect catalogue stays sealed while later experiments run blind against it. One run per model, so run-to-run variance is unmeasured. The graders hold the answer key and the runs do not, which is the point, but it means grading judgment calls exist; every one is recorded with its alternative reading in the grading file, and the two that could have moved a verdict were ruled explicitly.

What this does not prove: anything about live negotiation, about real retailer data feeds, or about models not tested. What it does establish, three times over on one dataset: a frontier agent can assemble a fact base that is overwhelmingly right and still hand a negotiator the one number that gets them caught, with full confidence, and the reflex that protects against it, doubting the counterparty's identity mappings, did not fire in any model tested.

NextThe register