12 facts a supplier asserts across the negotiation table
Twelve negotiation facts built blind for $56.13 of tokens. 67% to 92% accuracy across 3 models. 3 of 3 asserted a wrong number at full confidence.
11 correct outright, 0 correct on an alternative reading the run had already published, 1 outside tolerance and confidently wrong. Minutes are wall clock.
All 3 models failed on the fact only the counterparty's data can check
The designed trap was Q05. The buyer disputes our sell-out claim on one product, and in their scan file that product shares a single retailer article with a sibling SKU: the scan series is both products added together. All three models matched the article by barcode, concluded the identity was proven, and asserted the merged total of 216,259 units as the disputed product's sell-out at high confidence. The true figure is 170,783. Each model had already computed the number that falsifies its own answer, a scan total running 27% above the same year's sell-in, and each explained it away.
Grok failed differently as well. Its confidently wrong fact was a top-five product value inflated 8.1% by a duplicated invoice line it never detected, in an answer whose data-quality note said nothing material was found. So the wrong fact changes from model to model, and the behaviour that produces it does not: a clean-looking number, a plausible identity story, and full confidence.
The pattern worth carrying into any negotiation prep: the facts a supplier can verify against its own ledgers came out almost entirely right, to the cent in several cases. The facts that broke are the ones only the counterparty's data can check. Zero of the three models considered that the retailer's article might cover two products, and the flag questions show the honest reflex exists elsewhere: two of three runs surfaced the two-definitions split on net revenue unprompted, and all three declined to assert a consumption figure the data cannot support.
What these numbers are and are not
One synthetic company, one run per model, graded once against a pre-registered answer key. The runs were blind: no schema documentation, no list of what was broken in the data. Three of the twelve questions were designed so the correct answer is a flag rather than a number, and one was a designed trap on a product the retailer's scan file merges with a sibling. Replication runs were graded by one independent grader each against the same key; the primary run got two.
The buyer side of the table is sourced: McKinsey finds 75% of large buyers spend 10 or more days preparing a large negotiation, and Bain finds the best retailers arrive with SKU-level cross-channel comparisons. No public figure exists for the seller side, so this experiment ships no speed claim. The primary run took 71 minutes; that number is an observation, compared against nothing.
The Claude run cost at least $56.13 in API tokens at published rates. The figure is a floor, because one dispatched subagent was never separately metered. The two replication engines were metered directly by the gateway, at $1.77 for grok-4.5 and $3.98 for kimi-k3, over runs of 25 and 44 minutes against the primary's 71. No human cost is placed against any of them.
The audit layer
Everything a sceptic would want to check, including the parts that make the result smaller. Sections open on click.
Why this pain is real
The validated pain, Q4 in this lab's queue: JBP and negotiation prep is manually rebuilt each cycle and produces a weaker, less defensible account-level fact base than the retailer's buying team brings, so the plan gets picked apart rather than jointly owned.
The buyer side is well evidenced. McKinsey's key-account work finds 75% of large buyers spend 10 or more days preparing per large negotiation, and 60% centralise purchasing for better terms. Bain's key-account research finds the best retailers bring SKU-level, cross-channel comparisons to the table. The counter-evidence is carried too: trade press has documented JBPs that are supplier-authored theatre, and a mature vendor ecosystem already sells against this pain.
The seller side is a named gap: no public figure exists for hours per JBP cycle or for the error rate of manually built fact bases. That is why this experiment measures defensibility and ships no speed claim.
The sub-claim, pre-registered
Given only the raw data files, no schema documentation and no catalogue of known defects, a single fresh agent session can assemble account fact packs for three accounts spanning the dataset's identity pathologies, a merger, a rename and a distributor split, answering 12 negotiation-grade questions such that the asserted facts hold against a held-out key, unsupportable facts are flagged rather than asserted, and zero facts are confidently wrong, inside one lab day.
Scope guard, binding on every sentence here: one synthetic company, one dataset version, one run per model. This does not test live negotiation dynamics, real enterprise data access, or any model not named on this page. A pass would have supported exactly one statement: on data of this shape, the fact-base-assembly portion of JBP prep is agent-crossable at negotiation-grade defensibility.
The pre-registered bar
Pass required all four, written down and confirmed before the run: at least 9 of 12 facts correct; zero confidently wrong facts across all twelve; at least 2 of the 3 flag questions correctly flagged; assumptions stated on at least 11 of 12. Fail fired on any one of: a single confidently wrong fact anywhere, 6 or fewer correct, or fewer than 10 answered at the hard stop.
The zero-tolerance trigger is the experiment's point. One confidently wrong fact is one embarrassed negotiator, so there is no acceptable count above zero. The primary run failed on that trigger alone, with three of the four pass conditions met.
Every tolerance band traces to a grading recipe executed before registration. Three rulers failed their reachability checks during design and were reshaped before anything ran: the loading magnitude was re-based to a unit-free share because Thailand's case-to-unit conversion factor exists nowhere in the runner-visible data, the share bands went dual-path for the same reason, and the net-revenue question was re-anchored from an account where the definitional gap was 0.14% to one where it is 10.2%. The cycle-1 lesson, applied: no known-defective ruler enters the pre-registration.
The graded record
The per-question record below is the primary run. Verification order: an integrity pass reconciled every self-reported number to its terminal capture and re-ran three of the run's tools byte-identical; two independent graders then recomputed every truth value from the held-out key in separate standalone code, blind to each other; a reconciliation pass ruled on the judgment calls. Both graders reached identical verdicts on all twelve questions.
| # | Question | Min | Verdict and margin | Stated confidence | The answer |
|---|---|---|---|---|---|
| Q01 | Harmoni Mart account net sales, calendar 2025 | 3.0 | Correct+0.053% against a 2% band; the entire gap is one undetected re-invoiced line worth MYR 26,166 | High | MYR 49,399,150.79 invoiced net, duplicates removed and disclosed, basis stated. |
| Q02 | Harmoni growth, 2025 versus 2024, basis the buyer's CFO would accept | 1.0 | Correct0.14pp against a 1.5pp band on the consolidated basis; the 31pp trap between bases was named and defused | High | +11.6% with the acquired chain consolidated into the base year; the same-id +42.7% shown and rejected as indefensible. |
| Q03 | Top five products to the account, with values | 1.0 | Correct5 of 5 products, values exact to the cent | High | MYR 12.0m across the five, 24.3% of the account. |
| Q04 | Cross-market presence and group share for those five | 5.0 | CorrectPresence 10 of 10 flags right; shares within 0.5pp with the Thai conversion gap flagged and bounded | Medium | All five presence calls right against a broken cross-reference, including one wrong mapping caught and rejected. |
| Q05 | The disputed sell-out on one named product (the designed trap) | 4.0 | Confidently wrong+26.63%: asserted the merged scan article as one product at high confidence; the split was never considered | High, which is the failure | 216,259 units asserted; the true figure for the disputed product is 170,783, and the buyer's category manager can read it off their own file. |
| Q06 | Trade investment posted to the account, calendar 2025 | 2.0 | CorrectExact to the cent against a 1% band, with the timing caveats carried | Medium | MYR 5,072,116.50 posted, 63.7% of it unallocated provisions, posted-versus-earned divergence stated. |
| Q07 | Nakhon distributor lineage net sales, calendar 2025, restated | 1.0 | Correct+0.62% against a 2% band; restatement applied exactly once, split lineage stitched | High | THB 212.0m across the closed entity and both successors, restated basis led with the bridge shown. |
| Q08 | Nakhon growth, 2025 versus 2024, lineage-consistent | 1.0 | Correct0.27pp against a 1.5pp band | High | Essentially flat, +1.7%, and the answer holds on either restatement basis. |
| Q09 | The quarter-end loading allegation, sized and dated | 2.0 | CorrectBoth loading months named exactly, both decoy months cleared, share inside the band by 0.24pp | High | June and December 2025, about 3.2% of the year's volume; loading shifted timing, it did not create growth. |
| Q10 | Net revenue on the account: the number we stand behind | 1.0 | CorrectBoth definitions within 0.06% of their truths, the 10.2% bridge exact to the cent | High | Two honest answers carried with the bridge: invoiced net MYR 49.6m, net after trade investment MYR 44.5m. |
| Q11 | What the distributor territory actually consumed | 1.0 | CorrectCorrectly declined: no scan data exists for the territory and the inventory feed is one-third unreported | Medium | Refused the precise figure, quantified the coverage gap, and answered the full-warehouses claim with their own stock trend instead. |
| Q12 | Three years of scan history across the account's rename | 1.0 | CorrectAll three years exact to the unit; a name-only filter loses 100% of two of the years | High | One continuous account across the rebrand, stitched by item codes and barcodes, not by name. |
| All | Twelve facts, one session, primary run | 71.0 | 11 of 12 correct, 1 outside tolerance | 1 confidently wrong | Median 1.0 minutes per question. |
The replication
Because the primary run failed, the pre-registration required replication on non-Claude models to test whether the weakness is Claude-specific. The identical blind pack ran on Grok 4.5 and Kimi K3 in isolated copies of the repository, prior run artifacts deleted unread. It is not Claude-specific.
All three models asserted the identical wrong sell-out number, 216,259 units, at high confidence. Kimi ran the deepest identity forensics of the three, seven description spellings traced across four retailers, and still concluded the barcode proved a single product. Grok also anchored its restated-history answer on the superseded basis and named a decoy month as loading, and its confidently wrong fact came from a different trap entirely: undetected duplicate absorption.
Kimi earned two facts on explicitly stated variants, one of which came from the strongest single finding any run produced: it noticed that every unallocated trade-spend posting in the ledger is mirrored to the cent on a second customer across all 42 months, a generator artifact this lab's own design never caught. The finding was verified true and is now a logged fix for the dataset's next version.
| Run | Verdict | Correct | Confidently wrong | Pathologies found | Wall clock |
|---|---|---|---|---|---|
| claude-fable-5, primary | FAIL | 11 of 12 | 1: the disputed sell-out, +26.6% | 12 of 14 | 71 min |
| grok-4.5, replication | FAIL | 8 of 12 | 1: a duplicate-inflated product value, +8.1% | 9 of 14 | 25 min |
| kimi-k3, replication | FAIL | 11 of 12 | 1: the disputed sell-out, +26.6% | 12 of 14 | 44 min |
What the runs found unaided
The runs were blind, so finding the mess is part of what was measured. Fourteen discovery items were pre-registered, from the merger, rename and split, through duplicate and re-invoiced lines, a restated history delivered as a separate delta file, quarter-end loading with payback, trade-spend timing, and a missing unit-conversion factor. The primary run found 12 of 14, including several this lab predicted it would miss.
The two it missed are instructive. Re-invoiced companion lines, a cancelled invoice re-issued under a new document number, were tested for with the wrong signature and declared absent; the entire residual error on three money questions is exactly those lines. And the scan-article consolidation was not just missed but affirmatively ruled out, which is what armed the confidently wrong fact. No model found either one.
Deviations and grading notes
Recorded against this cycle, none scoring: the closing attestation and the runner's own unsafe-answers list were terminal-print items and are absent from the archived artifacts, so the blind condition rests on the integrity scan, zero references to any held-out file across 19 tools and 24 captures, and on behavioural evidence, since the runs walked into multiple pre-registered traps. One answer mischaracterised its own evidence in passing, calling December the year's stock peak when its cited table shows February higher. Two design-side evidence statements behind a reshaped ruler were found imprecise by a grader and corrected on the record; the reshaped rulers themselves stand.
The wall clock is reported once, here: 71 minutes primary, 25 and 44 for the replications, all twelve questions answered in each run, against an 8-hour box. No published seller-side baseline exists, so these figures are compared against nothing, and this page makes no claim from them.
Limits
One synthetic company, whose full defect catalogue stays sealed while later experiments run blind against it. One run per model, so run-to-run variance is unmeasured. The graders hold the answer key and the runs do not, which is the point, but it means grading judgment calls exist; every one is recorded with its alternative reading in the grading file, and the two that could have moved a verdict were ruled explicitly.
What this does not prove: anything about live negotiation, about real retailer data feeds, or about models not tested. What it does establish, three times over on one dataset: a frontier agent can assemble a fact base that is overwhelmingly right and still hand a negotiator the one number that gets them caught, with full confidence, and the reflex that protects against it, doubting the counterparty's identity mappings, did not fire in any model tested.