Control
AI capability or accuracy claims not substantiated by testing
Every public claim about what an AI feature can do, or how accurate it is, is backed by testing under the conditions the claim describes, recorded, and kept current.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
The guard to add
Back each published AI accuracy or capability claim with a recorded evaluation under the stated conditions, and show those conditions and limits next to the claim.
Consider a claims register in the repository (for example docs/claims-register.yaml) that maps every accuracy or capability statement in user-facing copy (landing pages, app strings, docs) to the evaluation run that supports it: task, dataset, population, date, model version, and result. The copy itself states the conditions and known limits beside the figure or links to the evaluation report. A CI step flags register entries whose model version no longer matches production so the claim is re-tested or withdrawn. Prefer copy that does not present the AI as a lawyer or financial adviser, or as a replacement for one.
Where it goes: 14 user-facing text, 12 repository artifacts, 13 tests and evals, 11 CI/CD pipeline.
What reviewers look for: for each percentage-accuracy or capability claim in user-facing copy, a register entry, model card, or eval report under the same task, data, and population, with conditions and limits shown or linked beside the claim; a re-test when the model changes; no 'AI lawyer' or 'replaces your financial adviser' copy.
Example (Next.js marketing page), before:
<p>Our AI invoice reader is 99% accurate.</p>After:
<p>
On 2,400 English invoices from US vendors (eval run 2026-08-14, model v4), our AI reader
extracted totals that matched human entry 96 times in 100. It is less reliable on handwritten
or non-English invoices. <a href="/evals/invoice-reader-2026-08">How we tested</a>
</p>Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Standard / soft law (1)
- International (INTL)
- AI actors should provide meaningful information that fosters understanding of AI systems' capabilities and limitations OECD/LEGAL/0449 — Principle 1.3
TwinEthos recommendation (not law) (1)
- Everywhere (*)
- Back every AI capability or accuracy claim with testing under the conditions it describes TwinEthos derivation — guardrail.ethics-substantiated-capability-claims · advisory
Related incidents
- FTC order bars Workado's unsubstantiated 98% AI-detector accuracy claim (2022-11; alleged (not proven)). The FTC's complaint alleges that Workado (formerly Content at Scale AI) advertised its AI Content Detector as 98% accurate based on the developers' published results for academic text from an open-source model it did not build or fine-tune, while promoting it for marketing and other non-academic text; the complaint says the developers' own data showed 53.2% accuracy on AI-generated non-academic text. Workado settled without admitting or denying the allegations; the final order (August 2025) bars accuracy or efficacy claims without competent and reliable evidence. Source: U.S. Federal Trade Commission (press release, 2025-08-28) · evidence grade: primary · cited by Back every AI capability or accuracy claim with testing under the conditions it describes
- FTC order bars DoNotPay's unsubstantiated 'robot lawyer' claims (2021; alleged (not proven)). The FTC's complaint alleges that DoNotPay marketed its subscription service as 'the world's first robot lawyer' without testing whether its law-related features performed like a human lawyer and without retaining attorneys to test their quality and accuracy. DoNotPay settled without admitting or denying the allegations; the final order (announced February 2025) requires $193,000 in monetary relief and notice to 2021-2023 subscribers, and bars claims that the service performs like a real lawyer without sufficient evidence. Source: U.S. Federal Trade Commission (press release, 2025-02-11) · evidence grade: primary · cited by Back every AI capability or accuracy claim with testing under the conditions it describes