Control
Safety or data screening of AI inputs or outputs applied to a sample instead of every item
Moderation, PII and data-loss screening, and other safety checks on AI inputs and outputs run on every item that reaches a person or a system of record; sampling is used only for human quality review on top of full screening, never in place of it.
Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.
Reach
Trust and provenance
How far the rules this guard addresses have been checked. Each rule links to its provision, with its citation, official text and its own panel.
- This control
- Audit-grade: meets all 3 checks of the TwinEthos audit standard that apply to it.
- Lanes
- TwinEthos recommendation (not law) 1
- Data release
- Data release 2026.10.03.4, data as of 3 Oct 2026, schema 0.3.10. This page also reflects corpus changes made after that release; they ship in the next one.
- Legal review
- None of the 1 rule has been reviewed by a lawyer; no TwinEthos rule has been legally reviewed yet. Treat each as research to check against the official text; it is not legal advice. Open questions for counsel on them: 1.
- Audit standard
- 1 of 1 rule audit-grade. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
- 1 detector, all experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify. Each provision lists its detectors' known limits.
The guard to add
Screen every AI input and output that reaches a person or record; sample only for extra human review.
Call the moderation, PII-screening or safety check on every input and output on the serving path, with no random.random() < rate, Math.random() or hash-modulo condition in front of it and no sample-rate setting below 1.0. If screening cost or latency is the problem, use a cheaper first-pass classifier on everything and escalate, cache verdicts for identical content, or batch, rather than skipping items. Keep random sampling for human quality-review queues that sit on top of the full automated check.
Where it goes: 1 application source code, 3 config and feature flags, 9 AI output handling.
What reviewers look for: guard calls that are unconditional on the serving path; no sampling condition or sample-rate configuration on moderation, PII or safety checks; sampling used only to route already-screened items to human review.
Example (Python + OpenAI moderation), before:
if random.random() < MODERATION_SAMPLE_RATE: # 0.1 to save cost
if moderate(reply):
return BLOCKED
return replyAfter:
if moderate(reply): # every reply
return BLOCKED
if random.random() < REVIEW_SAMPLE_RATE:
review_queue.put(reply) # sampling only for extra human review
return replyEngineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
TwinEthos recommendation (not law) (1)
- Everywhere (*)
- Screen every AI input and output; use sampling only for extra human review TwinEthos derivation — guardrail.opint-screen-every-item-not-a-sample
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.