Research

Commerce Index

What the agent is observing across the catalogue, published on a schedule.

The Commerce Index is a periodic account of what the FlowFinds agent is observed doing when it is asked to reason about commerce. It is a record, not a ranking. Each series below is computed at build time from the recorded run artifacts and the scenario source files; no figure on this page is written by hand, and a series for which no artifact exists is described rather than estimated.

This edition covers observations recorded up to 2026-09-05, drawn from 18 runs holding 51 scored cases, read from the live benchmark repository on the machine that built this page.

Status of these figures

All 18 recorded runs are filed under the status public-development-diagnostic-not-independent-benchmark. They are development diagnostics collected by FlowFinds Solutions on its own harness. They are not an independent benchmark, they were not administered by a third party, and they are not evidence of comparative model quality.

Provenance of this build

Every figure on this page was computed when the site was built, by reading 18 recorded runs containing 51 scored cases, from the benchmark repository present on the build machine. All of them were recorded on 2026-09-05. Nothing on this page is typed by hand; re-running the benchmark and rebuilding is the only way to change it.

Source files published alongside the runs: supplier-scenarios.ts, economics-scenarios.ts, claude-adapter.ts, eval-commerce.ts.

Methodology

The index has one rule, and every other property follows from it: a number appears here only if a file on disk produced it. The pipeline is the same one the rest of the research section uses. At build time it reads each run artifact — a JSON record written by the evaluation harness — and each scenario definition, aggregates them, and hands the result to this page. There is no database between the artifact and the sentence you are reading, and no editorial step in which a figure is chosen.

  1. Collection. Every scored case in every recorded run is included. Runs are not selected, excluded or reweighted, and failing runs enter the index on the same terms as passing ones.
  2. Derivation. Counts are counts. Rates are the count divided by the number of attempts recorded for that quantity, with the denominator stated beside it. Where no attempt is recorded the value is reported as not measured rather than as zero.
  3. Attribution. Each run pins the SHA-256 digest of the source files that produced it; 10 files are pinned across the recorded set, so a change to a scenario cannot silently change the meaning of an older observation.
  4. Publication. The index is recomputed on every build of this site. An edition is what a build produced, and it is dated by the most recent observation it contains rather than by the day it was published. When artifacts are added, the figures move on the next build without anyone editing this page.

What the index is measuring

The observations are of an agent answering fixed commerce prompts under a harness, not of a market. That is a narrower object than the name might suggest, and the narrowness is deliberate: it is the part of the system whose behaviour can be recorded exactly, replayed, and checked by a reader who disagrees with us. The scope of what is currently recorded is set out in the series below, and the scope of what is not is set out beneath them.

Series A — Composition of the observation corpus

What the agent is asked. 14 scenarios are currently declared across 3 capability groups, carrying 37 assertions in total — a mean of 2.6 per scenario. The composition matters because it determines what the index can see at all.

Property of the corpusCountShare of declared scenarios
Scenarios declared in source14100%
Requiring the agent to withhold a claim (a prohibition assertion)536%
Carrying a recorded reference answer857%
Seeding a prior turn before the prompt17%

The prohibition share is the one to read closely. It is the fraction of the corpus on which the agent is scored partly for what it declines to assert — a figure it cannot derive from the inputs, or an action it has not actually taken. A corpus without such cases would reward confident invention, so this share is a property of the measuring instrument and not a result.

Assertion kindOccurrencesShare of assertions
match3286%
doesNotMatch514%

Series B — What the agent reaches for

Across the recorded set, 51 of 51 scored cases — 100% — involved at least one tool call before the agent answered. No recorded case was answered from the prompt alone. The distribution below counts every call, and separately the number of cases in which each tool was called at least once, so that a tool called repeatedly within one case is not mistaken for a tool in wide use.

ToolCallsCases in which it was calledShare of scored cases
business5151100%
economics161631%
supplier101020%
revenue9918%
product_evidence7714%
store_funnel6612%
supplier_order6612%

Median time to produce one scored answer, across every case that recorded an elapsed time: 22.0 s. This is harness latency under the recorded conditions and is not a service-level statement about the product.

Series C — Computations the system declined

A refused tool call is not a model failure. It is a guard rail firing: the harness declining to compute on inputs it judges insufficient. These are published because they describe the boundary of what the system will assert, which is harder to see from successes than from refusals.

3 refusals recorded, across 2 tools.

ToolRefusalsReasons recorded, verbatim
economics2
  • Error: The supplied amounts mix tax-inclusive and tax-exclusive bases. Ask for amounts on one consistent tax basis; do not calculate profit or CPA from these values.
  • Error: landed_cost was not supplied in the quoted inputs; use null for an unknown cost.
supplier_order1
  • Error: Supply one complete, verbatim founder message containing the supplier terms.

Series D — How the agent fails, when it fails

Failures are classified by the error the harness recorded, not by inspection of the answer. The taxonomy is mechanical, so a failure cannot be filed under a gentler heading than the evidence supports.

ModeWhat the classification meansCasesShare of scored cases
Degraded output: no valid final answerThe agent loop returned a response the runner marked degraded, so no final message was available to score.1223.5%
Required content absentThe answer omitted a value, entity or qualification the scenario requires the model to state.12.0%
Prohibited content emittedThe answer contained a string the scenario forbids — typically a figure presented as established when the inputs do not support it.12.0%

Hardest and easiest observed cases

Ranked by pass rate over every recorded attempt. Attempts per case are small and unequal; these are the extremes of what has been observed, not a difficulty scale.

CaseAttemptsPassedPass rate
economics-06-tax-basis4250%
supplier-expired-quote2150%
unit-economics8450%
missing-costs5360%
conversion-diagnosis4375%
economics-09-order-budget22100%
supplier-deadline22100%
supplier-missing-landed22100%

Series E — Coverage of the index itself

An index that reported only what it had measured, without reporting what it had not, would be misleading by omission. This series is the index measuring its own gaps.

Coverage propertyValueWhat it means for the figures above
Declared scenarios with no recorded attempt0These capabilities are defined but unobserved. Every series above is silent about them.
Recorded cases no longer declared in source0Historical observations whose definition has since changed or been withdrawn. They are counted, and they are not re-runnable as recorded.
Runs whose artifact omits the provider field3These cases cannot be attributed to a provider and are grouped under an unrecorded provider rather than assigned to one.
Distinct provider and model pairings observed4Sample sizes differ between pairings, so no comparison between them is offered here.
Observation window2026-09-05Nothing outside this window is described. A short window cannot show a trend, and none is claimed.

Series defined but not yet published

Two series belong in an index of this name and are absent from this edition. They are named here, with the condition each one requires, so that their absence is legible rather than convenient.

When either series is published it will appear here computed from artifacts on the same terms as the series above, with its sampling frame stated. It will not appear as an estimate in the meantime.

Limitations

Checking this edition

Every artifact behind these figures is downloadable in full from the results page, the protocol that produced them is written out on the methodology page, and the commands that regenerate them are on the reproduction page. If a figure on this page is wrong, those three pages are sufficient to demonstrate it. If you would rather learn the reasoning the index is testing than read the index, it is taught in full and for free in learn.