Recommended guardrail
Screen every AI input and output; use sampling only for extra human review
Run moderation, PII screening and other safety checks on every AI input and output that reaches a person or a system of record, and keep random sampling for human quality review on top of that, never in place of it. Detect guard calls gated by random or hash-based sampling, and screening sample-rate settings below 1.0.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.
Trust and provenance
- Lane
- TwinEthos recommendation (not law) TwinEthos recommendation, not law
- Official source
- TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
- Data release
- Data release 2026.10.03.4, data as of 3 Oct 2026, schema 0.3.10. This page also reflects corpus changes made after that release; they ship in the next one.
- Legal review
- Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet. Open questions for counsel on this rule: 1.
- Audit standard
- Audit-grade: meets all 11 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
1 detector (code pattern), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.
Known limits:
- Sampling that sends already-screened items to a human review queue is the recommended pattern; it is not reported because review calls are not guard calls.
Evidence grade
Related law in force in 1 jurisdiction · 2 standards and frameworks · 0 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law on this control itself is in force in the corpus; binding law in provisions cited as convergence, which cover part of the control or a related one, is in force in 1 jurisdiction (IN). 2 standards and frameworks recommend it (MITRE ATLAS, NIST GenAI Profile (AI 600-1)). 0 graded incidents cited.
Related law in force (cited as convergence; not on this control itself)
- Generation tools must deploy technical measures that stop users creating unlawful synthetic content (India IT Rules 2026) (India (IN); IT Rules, 2021, rule 3(3)(a)(i) (due diligence in relation to synthetically generated information: no unlawful synthetic content), inserted by G.S.R. 120(E); cited)
Standards and frameworks
- GenAI applications should filter inputs and outputs for harmful, illegal, or violent content (NIST GenAI Profile MG-3.2-005) (NIST GenAI Profile (AI 600-1); NIST AI 600-1, MG-3.2-005 (content filters on GAI inputs and outputs); cited)
- GenAI outputs should be screened for PII/sensitive-data leakage (NIST GenAI Profile) (NIST GenAI Profile (AI 600-1); NIST AI 600-1, Sec. 2 risk list (Data Privacy); cited)
Published AI security standards mapping to this control
- MITRE ATLAS AML.M0020: Generative AI Guardrails · crosswalk status: covered
Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.
Family “AI controls are not preserved under cost, latency, or model-change pressure”: binding law on related controls is in force in Illinois (US-IL). Context only: it does not change this guardrail's grade.
The guard to add
Screen every AI input and output that reaches a person or record; sample only for extra human review.
Call the moderation, PII-screening or safety check on every input and output on the serving path, with no random.random() < rate, Math.random() or hash-modulo condition in front of it and no sample-rate setting below 1.0. If screening cost or latency is the problem, use a cheaper first-pass classifier on everything and escalate, cache verdicts for identical content, or batch, rather than skipping items. Keep random sampling for human quality-review queues that sit on top of the full automated check.
Example (Python + OpenAI moderation), before:
if random.random() < MODERATION_SAMPLE_RATE: # 0.1 to save cost
if moderate(reply):
return BLOCKED
return replyAfter:
if moderate(reply): # every reply
return BLOCKED
if random.random() < REVIEW_SAMPLE_RATE:
review_queue.put(reply) # sampling only for extra human review
return replyControl: Safety or data screening of AI inputs or outputs applied to a sample instead of every item. Engineering guidance, not legal advice.
Why
Screening is easy to thin out when cost or latency bites, and sampling looks responsible on a dashboard: the check runs, flags appear, the numbers seem plausible. But a guard that runs on a tenth of outputs lets most harmful or leaking outputs through, and the users who receive them are not a sample. NIST's Generative AI Profile recommends filtering inputs and outputs for harmful content and screening outputs for personal data, and MITRE ATLAS describes moderation of prompts and responses before they reach users; this guardrail applies those recommendations to every item.
Class: operational integrity · set: operational integrity · maturity: reviewed · confidence: medium · id guardrail.opint-screen-every-item-not-a-sample
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.