commerce-v1

commerce-v1 results

Every recorded run, head to head, with dates and model versions.

Status of these figures

All 18 recorded runs are filed under the status public-development-diagnostic-not-independent-benchmark. They are development diagnostics collected by FlowFinds Solutions on its own harness. They are not an independent benchmark, they were not administered by a third party, and they are not evidence of comparative model quality.

Provenance of this build

Every figure on this page was computed when the site was built, by reading 18 recorded runs containing 51 scored cases, from the benchmark repository present on the build machine. All of them were recorded on 2026-09-05. Nothing on this page is typed by hand; re-running the benchmark and rebuilding is the only way to change it.

Source files published alongside the runs: supplier-scenarios.ts, economics-scenarios.ts, claude-adapter.ts, eval-commerce.ts.

How to read this page

Each table below is computed from the recorded artifacts at build time. Runs differ in which scenarios they contain and in which revision of the harness produced them, so a pass rate here is the proportion of attempted cases that passed — not a normalised score, and not comparable across models without first checking that the same scenarios and the same pinned sources were involved. The methodology page sets out what is and is not controlled for.

Across 18 runs there are 51 scored cases, of which 37 passed and 14 did not. 12 cases produced no valid final answer and are counted among the failures.

Head to head

Grouped by provider and requested model. Ordered by number of scored cases, which is not a ranking.
ProviderRequested modelRunsCasesPassedPass rateDegradedMedian case timeFirst runLast runResolved model identifiersRecorded cost
openaigpt-5.57181478%223.6 s2026-09-052026-09-05not reportednot recorded
anthropicclaude-fable-5-151616100%022.1 s2026-09-052026-09-05claude-fable-5-1 claude-haiku-4-5-20251001 USD 3.06
unrecordedgpt-5.5314643%817.2 s2026-09-052026-09-05not reportednot recorded
anthropicfable33133%21.3 s2026-09-052026-09-05claude-fable-5-1 claude-haiku-4-5-20251001 USD 0.23

The resolved-identifier column matters: a requested model name and the identifier that actually served the request are recorded separately, and where they differ the difference is shown rather than smoothed over. Recorded cost is the adapter’s own estimate and is only present for providers that report it, so it is not a cost comparison.

Per-scenario breakdown

Cells read passed / attempts across every recorded run. An empty cell means the model was never put to that scenario.
ScenarioAttemptsPassedPass rateopenai · gpt-5.5anthropic · claude-fable-5-1unrecorded · gpt-5.5anthropic · fable
unit-economics8450%1 / 11 / 11 / 31 / 3
conversion-diagnosis4375%1 / 11 / 11 / 2
missing-costs5360%1 / 11 / 11 / 3
memory-budget4375%1 / 11 / 11 / 2
mixed-currency4375%1 / 11 / 11 / 2
supplier-quote4375%1 / 11 / 11 / 2
economics-04-loss22100%1 / 11 / 1
economics-06-tax-basis4250%1 / 31 / 1
economics-07-decimal-comma22100%1 / 11 / 1
economics-09-order-budget22100%1 / 11 / 1
supplier-minimum-order6583%2 / 33 / 3
supplier-deadline22100%1 / 11 / 1
supplier-missing-landed22100%1 / 11 / 1
supplier-expired-quote2150%0 / 11 / 1

9 scenarios have recorded at least one failure: unit-economics (4/8), conversion-diagnosis (3/4), missing-costs (3/5), memory-budget (3/4), mixed-currency (3/4), supplier-quote (3/4), economics-06-tax-basis (2/4), supplier-minimum-order (5/6), supplier-expired-quote (1/2). The failure text for each is in the raw artifact, and the modes they fall into are analysed in the system cards.

Every recorded run

One row per recorded run, newest first. Each identifier is the run directory the harness created.
RunMeasured at (UTC)ProviderModelCasesPassedPass rateExpected casesPinned filesRaw artifact
diagnostic-FU1SXM2026-09-05 13:38:33Zopenaigpt-5.511100%1, complete10diagnostic-FU1SXM.json
diagnostic-tGXwE72026-09-05 13:37:34Zopenaigpt-5.5100%1, complete10diagnostic-tGXwE7.json
diagnostic-t3D8Tl2026-09-05 13:36:26Zanthropicclaude-fable-5-144100%4, complete10diagnostic-t3D8Tl.json
diagnostic-nwH3rH2026-09-05 13:36:26Zopenaigpt-5.54375%4, complete10diagnostic-nwH3rH.json
diagnostic-1zdMED2026-09-05 13:10:02Zanthropicclaude-fable-5-111100%1, complete9diagnostic-1zdMED.json
diagnostic-3HKjOI2026-09-05 13:09:52Zopenaigpt-5.511100%1, complete9diagnostic-3HKjOI.json
diagnostic-vDO1md2026-09-05 13:08:48Zanthropicclaude-fable-5-111100%1, complete9diagnostic-vDO1md.json
diagnostic-JDDdTU2026-09-05 13:08:44Zopenaigpt-5.5100%1, complete9diagnostic-JDDdTU.json
diagnostic-Bl0Qup2026-09-05 13:04:33Zopenaigpt-5.54375%4, complete7diagnostic-Bl0Qup.json
diagnostic-Cpb4yz2026-09-05 13:04:01Zanthropicclaude-fable-5-144100%4, complete7diagnostic-Cpb4yz.json
diagnostic-DI21Oc2026-09-05 12:53:15Zanthropicclaude-fable-5-166100%6, complete6diagnostic-DI21Oc.json
diagnostic-wq3uL52026-09-05 12:53:13Zopenaigpt-5.566100%6, complete6diagnostic-wq3uL5.json
diagnostic-xeNGjl2026-09-05 12:50:01Zanthropicfable11100%1, complete6diagnostic-xeNGjl.json
diagnostic-EebtBs2026-09-05 12:48:30Zanthropicfable100%1, complete6diagnostic-EebtBs.json
diagnostic-v3BetW2026-09-05 12:48:01Zanthropicfable100%1, complete5diagnostic-v3BetW.json
diagnostic-f5udhY2026-09-05 12:43:36Zunrecordedgpt-5.522100%2, complete5diagnostic-f5udhY.json
diagnostic-5oHnnI2026-09-05 12:43:30Zunrecordedgpt-5.56467%6, complete5diagnostic-5oHnnI.json
diagnostic-XxtsQO2026-09-05 12:40:37Zunrecordedgpt-5.5600%6, complete5diagnostic-XxtsQO.json

Raw artifacts

18 run artifacts of 18 are published verbatim under /research/commerce-v1/runs/, one JSON file per run. Each contains the measurement timestamp, the requested model and provider, the source digests, the adapter metadata, and for every case the assertions’ outcome, the tools called, the tool diagnostics, the elapsed time, the full response and, where the case failed, the failure text. They are the same files the tables above are computed from.

This build read its figures from the benchmark repository on the build machine. Any run recorded since the published snapshot was taken appears in the tables with no download link, marked not yet published, rather than being hidden.

Model pairings measured so far: openai gpt-5.5, 18 cases, 78% passed, last run 2026-09-05; anthropic claude-fable-5-1, 16 cases, 100% passed, last run 2026-09-05; provider unrecorded gpt-5.5, 14 cases, 43% passed, last run 2026-09-05; anthropic fable, 3 cases, 33% passed, last run 2026-09-05.

Read next

To check these numbers rather than take them, follow the reproduction page: it lists the commands that produced every artifact above. For what the failures were, read the system cards.