Research

Why we built the benchmark before we built the demo

commerce-v1 was fixed — scenarios, rubric and assertions — before a single run was recorded against it.

A commerce agent is easy to demonstrate and hard to evaluate. It writes fluent copy, it produces confident margin arithmetic, and none of that tells a founder whether it will be right about the thing that costs them money. So we wrote the measurement first. commerce-v1 declares its scenarios in source files, groups them by capability, and scores each one with assertions over the agent's final message. The suite existed and was frozen before we had a number to report against it.

What it measures is deliberately narrow: the reasoning a founder needs before money moves. Arithmetic on quoted terms. The difference between a price and a landed cost. The difference between revenue and profit. The difference between a written quote and a secured commitment. And the discipline to answer that something cannot be determined from the inputs given, when that is the truthful answer.

That last one is scored as a capability rather than treated as a caveat. Several scenarios are graded partly on what the final message must not contain — a figure asserted as established when it was only quoted, or a claim that an order has been placed when nothing was placed. An agent that refuses to overstate scores for it. An agent that fills a gap with a plausible number loses the case outright.

The whole definition is published, including the parts that make us look worse: the run protocol, the source hashing, and the recorded attempts that failed.

Read next

Read the commerce-v1 methodology.

More from FlowFinds

To fact-check anything above before you publish it, write to [email protected].