Trust

Safety approach

The limits we put on agent authority, and why.

An agent that runs a shop is asked to do things that cannot be undone: spend money, publish a claim about a product, answer a buyer who will act on the answer. Our approach to safety is mostly a set of powers we have declined to give it, and this page is the list, with the reasoning attached to each one.

The principle

Authority is granted in proportion to how reversible the action is. A draft can be rewritten, so the agent writes freely. A published post cannot be unpublished from everyone who saw it, so the agent does not publish. A payment cannot be recalled, so the agent does not pay. This is not caution about model quality in the abstract; it is a judgement about which mistakes a person can absorb and which they cannot.

What the agent is not allowed to do

What it is allowed to do, and why that is safe enough

The agent researches products, prices an offer against listings it has observed, generates and regenerates a storefront, drafts advertising with pass criteria attached, and monitors orders and satisfaction. Each of those produces an artefact a person reviews before it has any effect on the world, and each is shown with the evidence it was derived from — the source row behind a price, the reason behind a section order, the listings behind a margin claim. An artefact you can trace is one you can reject.

Grounding: evidence, not confidence

The product is built so that a claim carries its source. A find is presented with the observations that support it; a guided edit records why it was made and what it was based on; the dashboard’s verdict is a summary of states the system can point at rather than a mood. The failure mode we care most about is a fluent, unsupported assertion, and showing the source row is the cheapest defence against it.

Honest limits

The agent will be confidently wrong sometimes. It can misread a market, propose a price that does not hold, or write copy that overstates a product. It has no way of knowing whether a supplier will actually ship. It cannot tell whether a claim about a product is legally permissible in the market you sell into — that judgement is yours, and the acceptable use policy is explicit that generated wording is never a defence for a misleading claim.

Rather than assert that these failures are rare, we measure them. The benchmark, its scoring rubric and the runs the agent failed are published at research, and the per-model limitations at system cards.

Where we draw a line on use

Some businesses we will not help run: counterfeits, unlicensed regulated goods, deceptive selling and schemes that pay for recruitment rather than sales. The full boundary is the acceptable use policy, and enforcement is proportionate, reviewed by a person, and appealable to [email protected].

Telling us when it fails

A harmful or badly wrong output is worth reporting to [email protected]; a way to make the agent do something it should not goes to [email protected] under the responsible disclosure policy. Both change the product; the second is how the limits above get tightened.