Company

Mission

Why a commerce agent should be measured like a frontier model, not sold like a template.

Two industries sell software that starts an online shop. One sells templates: you get a theme, a checkout and a blank catalogue, and every judgement that decides whether the shop works is still yours. The other sells an agent: the judgements are made for you, by a model, and the quality of those judgements is the entire product.

Those two things are sold in the same way — features, screenshots, a price — and they should not be. A template can be inspected before purchase. An agent cannot. What you are buying is a distribution of decisions you will not see until it has already made them.

The argument

When a laboratory ships a language model, it publishes an evaluation suite, the score on each task, the conditions of the run, and a description of where the model fails. Nobody considers this generous. It is the minimum required for the claim to mean anything, because the capability is not visible from the outside.

A commerce agent has exactly that property. “It finds winning products” is not a feature; it is a claim about a success rate under conditions that are never stated. The honest version of the sentence has a number, a task definition and a date attached. If we cannot attach those, we should not say it.

So: a commerce agent should be measured like a frontier model, not sold like a template. That is the whole mission. Everything below is what follows from taking it seriously rather than saying it.

What follows from it

1. A fixed public benchmark, or the claim does not ship

A capability we cannot state as a score on a defined task does not become a marketing sentence. commerce-v1 defines the tasks: how a scenario is constructed, what a correct answer looks like, how a run is scored, and how the inputs are hashed so a run can be identified later. The definition is published before the results, which is the only order in which a benchmark constrains anyone. Read the methodology.

2. We publish the failures

A results page containing only successes is an advertisement wearing a lab coat. Every recorded run appears in the results, scored, dated and attributed to a model version — including runs where the agent chose a product it should not have, priced an offer it could not defend, or produced a plan that did not survive contact with a supplier. Failure modes are the most useful thing we know about our own system, and hiding them mostly deceives us. See every run.

3. We publish the methodology, so you can disagree with it

A benchmark built by the company it flatters is worth what its rubric is worth. Ours is written down in enough detail to be attacked: scenario sources, the rubric each dimension is scored against, the run protocol, and the procedure for adding or retiring a scenario. If the rubric is wrong, that is a finding we want, and it is cheaper for us to receive it than to discover it from a customer. Run it yourself.

4. A system card for every model we ship

Each model that runs a customer’s store gets a card: what it is intended to do, what it was evaluated on, the limitations we know about, the failure modes we have observed, and the authority it has been granted over money and communications. When we replace a model, the card is replaced with it and the previous one stays published, because a customer who made a decision under the old card is entitled to see it. Read the system cards.

5. Numbers live in one place

Every figure on this site is read from a recorded run at build time and rendered in the research section. No other page quotes a score, because a number copied into prose is a number that will be stale within a release and true only in an archive. Pages that want to make a quantitative claim link to the run instead.

What this costs us

It is slower. A capability spends time in evaluation before it becomes a sentence anyone is allowed to write, and some capabilities do not survive that. It is also less flattering: a published failure is permanent, and competitors who publish nothing will always appear to have no failures at all. We think this is the correct trade, because the alternative is a market in which no claim about an agent can be checked and therefore no claim is worth anything, including ours.

The measurement is not a marketing asset that happens to sit next to the product. It is the constraint the product is built under. How we build describes the mechanism that enforces it.