commerce-v1
commerce-v1 results
Every recorded run, head to head, with dates and model versions.
Status of these figures
All 18 recorded runs are filed under the status public-development-diagnostic-not-independent-benchmark. They are development diagnostics collected by FlowFinds Solutions on its own harness. They are not an independent benchmark, they were not administered by a third party, and they are not evidence of comparative model quality.
Provenance of this build
Every figure on this page was computed when the site was built, by reading 18 recorded runs containing 51 scored cases, from the benchmark repository present on the build machine. All of them were recorded on 2026-09-05. Nothing on this page is typed by hand; re-running the benchmark and rebuilding is the only way to change it.
Source files published alongside the runs: supplier-scenarios.ts, economics-scenarios.ts, claude-adapter.ts, eval-commerce.ts.
How to read this page
Each table below is computed from the recorded artifacts at build time. Runs differ in which scenarios they contain and in which revision of the harness produced them, so a pass rate here is the proportion of attempted cases that passed — not a normalised score, and not comparable across models without first checking that the same scenarios and the same pinned sources were involved. The methodology page sets out what is and is not controlled for.
Across 18 runs there are 51 scored cases, of which 37 passed and 14 did not. 12 cases produced no valid final answer and are counted among the failures.
Head to head
| Provider | Requested model | Runs | Cases | Passed | Pass rate | Degraded | Median case time | First run | Last run | Resolved model identifiers | Recorded cost |
|---|---|---|---|---|---|---|---|---|---|---|---|
| openai | gpt-5.5 | 7 | 18 | 14 | 78% | 2 | 23.6 s | 2026-09-05 | 2026-09-05 | not reported | not recorded |
| anthropic | claude-fable-5-1 | 5 | 16 | 16 | 100% | 0 | 22.1 s | 2026-09-05 | 2026-09-05 | claude-fable-5-1 claude-haiku-4-5-20251001 | USD 3.06 |
| unrecorded | gpt-5.5 | 3 | 14 | 6 | 43% | 8 | 17.2 s | 2026-09-05 | 2026-09-05 | not reported | not recorded |
| anthropic | fable | 3 | 3 | 1 | 33% | 2 | 1.3 s | 2026-09-05 | 2026-09-05 | claude-fable-5-1 claude-haiku-4-5-20251001 | USD 0.23 |
The resolved-identifier column matters: a requested model name and the identifier that actually served the request are recorded separately, and where they differ the difference is shown rather than smoothed over. Recorded cost is the adapter’s own estimate and is only present for providers that report it, so it is not a cost comparison.
Per-scenario breakdown
| Scenario | Attempts | Passed | Pass rate | openai · gpt-5.5 | anthropic · claude-fable-5-1 | unrecorded · gpt-5.5 | anthropic · fable |
|---|---|---|---|---|---|---|---|
unit-economics | 8 | 4 | 50% | 1 / 1 | 1 / 1 | 1 / 3 | 1 / 3 |
conversion-diagnosis | 4 | 3 | 75% | 1 / 1 | 1 / 1 | 1 / 2 | — |
missing-costs | 5 | 3 | 60% | 1 / 1 | 1 / 1 | 1 / 3 | — |
memory-budget | 4 | 3 | 75% | 1 / 1 | 1 / 1 | 1 / 2 | — |
mixed-currency | 4 | 3 | 75% | 1 / 1 | 1 / 1 | 1 / 2 | — |
supplier-quote | 4 | 3 | 75% | 1 / 1 | 1 / 1 | 1 / 2 | — |
economics-04-loss | 2 | 2 | 100% | 1 / 1 | 1 / 1 | — | — |
economics-06-tax-basis | 4 | 2 | 50% | 1 / 3 | 1 / 1 | — | — |
economics-07-decimal-comma | 2 | 2 | 100% | 1 / 1 | 1 / 1 | — | — |
economics-09-order-budget | 2 | 2 | 100% | 1 / 1 | 1 / 1 | — | — |
supplier-minimum-order | 6 | 5 | 83% | 2 / 3 | 3 / 3 | — | — |
supplier-deadline | 2 | 2 | 100% | 1 / 1 | 1 / 1 | — | — |
supplier-missing-landed | 2 | 2 | 100% | 1 / 1 | 1 / 1 | — | — |
supplier-expired-quote | 2 | 1 | 50% | 0 / 1 | 1 / 1 | — | — |
9 scenarios have recorded at least one failure: unit-economics (4/8), conversion-diagnosis (3/4), missing-costs (3/5), memory-budget (3/4), mixed-currency (3/4), supplier-quote (3/4), economics-06-tax-basis (2/4), supplier-minimum-order (5/6), supplier-expired-quote (1/2). The failure text for each is in the raw artifact, and the modes they fall into are analysed in the system cards.
Every recorded run
| Run | Measured at (UTC) | Provider | Model | Cases | Passed | Pass rate | Expected cases | Pinned files | Raw artifact |
|---|---|---|---|---|---|---|---|---|---|
diagnostic-FU1SXM | 2026-09-05 13:38:33Z | openai | gpt-5.5 | 1 | 1 | 100% | 1, complete | 10 | diagnostic-FU1SXM.json |
diagnostic-tGXwE7 | 2026-09-05 13:37:34Z | openai | gpt-5.5 | 1 | 0 | 0% | 1, complete | 10 | diagnostic-tGXwE7.json |
diagnostic-t3D8Tl | 2026-09-05 13:36:26Z | anthropic | claude-fable-5-1 | 4 | 4 | 100% | 4, complete | 10 | diagnostic-t3D8Tl.json |
diagnostic-nwH3rH | 2026-09-05 13:36:26Z | openai | gpt-5.5 | 4 | 3 | 75% | 4, complete | 10 | diagnostic-nwH3rH.json |
diagnostic-1zdMED | 2026-09-05 13:10:02Z | anthropic | claude-fable-5-1 | 1 | 1 | 100% | 1, complete | 9 | diagnostic-1zdMED.json |
diagnostic-3HKjOI | 2026-09-05 13:09:52Z | openai | gpt-5.5 | 1 | 1 | 100% | 1, complete | 9 | diagnostic-3HKjOI.json |
diagnostic-vDO1md | 2026-09-05 13:08:48Z | anthropic | claude-fable-5-1 | 1 | 1 | 100% | 1, complete | 9 | diagnostic-vDO1md.json |
diagnostic-JDDdTU | 2026-09-05 13:08:44Z | openai | gpt-5.5 | 1 | 0 | 0% | 1, complete | 9 | diagnostic-JDDdTU.json |
diagnostic-Bl0Qup | 2026-09-05 13:04:33Z | openai | gpt-5.5 | 4 | 3 | 75% | 4, complete | 7 | diagnostic-Bl0Qup.json |
diagnostic-Cpb4yz | 2026-09-05 13:04:01Z | anthropic | claude-fable-5-1 | 4 | 4 | 100% | 4, complete | 7 | diagnostic-Cpb4yz.json |
diagnostic-DI21Oc | 2026-09-05 12:53:15Z | anthropic | claude-fable-5-1 | 6 | 6 | 100% | 6, complete | 6 | diagnostic-DI21Oc.json |
diagnostic-wq3uL5 | 2026-09-05 12:53:13Z | openai | gpt-5.5 | 6 | 6 | 100% | 6, complete | 6 | diagnostic-wq3uL5.json |
diagnostic-xeNGjl | 2026-09-05 12:50:01Z | anthropic | fable | 1 | 1 | 100% | 1, complete | 6 | diagnostic-xeNGjl.json |
diagnostic-EebtBs | 2026-09-05 12:48:30Z | anthropic | fable | 1 | 0 | 0% | 1, complete | 6 | diagnostic-EebtBs.json |
diagnostic-v3BetW | 2026-09-05 12:48:01Z | anthropic | fable | 1 | 0 | 0% | 1, complete | 5 | diagnostic-v3BetW.json |
diagnostic-f5udhY | 2026-09-05 12:43:36Z | unrecorded | gpt-5.5 | 2 | 2 | 100% | 2, complete | 5 | diagnostic-f5udhY.json |
diagnostic-5oHnnI | 2026-09-05 12:43:30Z | unrecorded | gpt-5.5 | 6 | 4 | 67% | 6, complete | 5 | diagnostic-5oHnnI.json |
diagnostic-XxtsQO | 2026-09-05 12:40:37Z | unrecorded | gpt-5.5 | 6 | 0 | 0% | 6, complete | 5 | diagnostic-XxtsQO.json |
Raw artifacts
18 run artifacts of 18 are published verbatim under /research/commerce-v1/runs/, one JSON file per run. Each contains the measurement timestamp, the requested model and provider, the source digests, the adapter metadata, and for every case the assertions’ outcome, the tools called, the tool diagnostics, the elapsed time, the full response and, where the case failed, the failure text. They are the same files the tables above are computed from.
This build read its figures from the benchmark repository on the build machine. Any run recorded since the published snapshot was taken appears in the tables with no download link, marked not yet published, rather than being hidden.
Model pairings measured so far: openai gpt-5.5, 18 cases, 78% passed, last run 2026-09-05; anthropic claude-fable-5-1, 16 cases, 100% passed, last run 2026-09-05; provider unrecorded gpt-5.5, 14 cases, 43% passed, last run 2026-09-05; anthropic fable, 3 cases, 33% passed, last run 2026-09-05.
Read next
To check these numbers rather than take them, follow the reproduction page: it lists the commands that produced every artifact above. For what the failures were, read the system cards.