Trust
Availability commitment
Our incident policy, what is protected first, and how outages are reported.
We do not promise an uptime percentage. The status page shows the one we measure — OK checks over all checks in the last ninety days, from our own probe, with the method printed beside the number — but a measured figure is not a contractual one, and we have not been operating long enough to offer a service level backed by credits. What follows is what we do commit to, which is checkable.
What we commit to
- A public status page. Current state, ninety days of history and every detected outage are on the status page, computed from a probe that checks every service every five minutes. It runs as a separate process from the product; if it is unreachable too, assume the whole host is down and write to support.
- Acknowledgement of a customer-visible incident within 30 minutes of confirming it, on the status page, with what is affected and what is not.
- Updates at least hourly while an incident is open, even when the update is that we are still working on it.
- A written post-incident review within five working days for any incident that caused data loss or more than an hour of unavailability, published on the status page: what happened, why, what we changed.
- Notice of planned maintenance at least 48 hours in advance, scheduled outside European business hours where the work allows it.
What is protected first
Not everything degrades equally. When capacity is constrained, we protect, in this order: the integrity of your data; the money paths, so that a payment is never lost or double counted; the ability to sign in and see your work; then generation and research, which can queue. An agent run that fails is an inconvenience. A payment that is credited twice is a financial error, and the system is designed so that a retried confirmation cannot cause one.
How an incident runs
- Detect. Automated checks and error reporting page the on-call engineer; so does a customer report, which is treated as a detection rather than a complaint.
- Declare. One person owns the incident. It goes on the status page as soon as it is confirmed to be customer-visible.
- Mitigate. Restore service first, understand it fully second.
- Communicate. Hourly updates, and a closing note saying the incident is resolved and what to do if you are still affected.
- Review. A written review, and — where the failure was one a test could have caught — a regression test added before the review is closed.
Data durability
Encrypted snapshots are taken regularly and retained on the rotation described at data handling and retention, which is also why a deleted record takes up to 30 days to disappear from backups. Restores are exercised rather than assumed; a backup nobody has restored is a hypothesis.
What is excluded
Features labelled beta or preview are excluded from the commitments above, as clause 7 of the service-specific terms states. So are outages caused by a third party outside our control — a payment provider, a network provider, a model provider — although we will report them on the status page and say who is affected, because knowing why you cannot work matters even when the cause is not ours.
Service credits
We do not currently offer a contractual uptime service level with credits. Where an outage made a paid plan unusable for a significant part of a billing period, we refund that part under clause 7 of the billing and refund terms. If your business requires a formal service level agreement, write to [email protected] and we will tell you honestly whether we can offer one yet.
Reporting an outage
Check the status page first, then write to [email protected] with what you were doing, the time and any error text. A security issue goes to [email protected] instead, under the responsible disclosure policy.