TwinEthos home

Recommended guardrail

Screen user input and generated output with a content-safety filter — everywhere

On every user-facing generation path, screen user input before the model call and generated output before it is shown, stored or sent, with a moderation endpoint, a safety classifier or a guardrail service whose categories match what the product can foreseeably produce; act on the result (block, redact, safe fallback or review) rather than only logging it; keep provider and pipeline safety filters switched on; fail closed when the filter errors; screen every item, not a sample; and cover every language the product serves. Apply this regardless of jurisdiction. Detect generation paths with safety filters switched off and moderation results that are never acted on.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.

Trust and provenance

Lane
TwinEthos recommendation (not law) TwinEthos recommendation, not law
Official source
TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
Data release
Data release 2026.10.05, data as of 5 Oct 2026, schema 0.3.11. This page also reflects corpus changes made after that release; they ship in the next one.
Legal review
Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet.
Audit standard
Audit-grade: meets all 11 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
Detectors

3 detectors (code pattern, data flow), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.

Known limits:

  • Filters disabled in a provider console (Azure OpenAI content-filter policies) rather than in code
  • A disabled provider filter replaced by an equivalent filter in the application (a moderation call on the same path) is acceptable; check for it. A threshold key set to OFF in non-AI configuration (for example a logging…
  • Provider-side filters configured outside the repository

3 more known limits in the data release.

Evidence grade

Law in force in 2 jurisdictions

Law in force in 2 jurisdictions · 4 standards and frameworks · 0 graded incidents.

TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control is in force in 2 jurisdictions (AU, IN). 4 standards and frameworks recommend it (MITRE ATLAS, NIST AI 100-2, NIST GenAI Profile (AI 600-1), OWASP LLM 2026). 0 graded incidents cited.

Law in force on this control or cited as convergence

Standards and frameworks

Published AI security standards mapping to this control

  • OWASP LLM 2026 LLM08: Hidden Context Exposure · crosswalk status: covered
  • MITRE ATLAS AML.M0020: Generative AI Guardrails · crosswalk status: covered
  • NIST AI 600-1 2.3: Dangerous, Violent, or Hateful Content · crosswalk status: covered
  • NIST AI 600-1 2.11: Obscene, Degrading, and/or Abusive Content · crosswalk status: covered
  • NIST AI 600-1 MS-2.6-006: Verify the system handles queries that may give rise to inappropriate, malicious, or illegal usage · crosswalk status: covered
  • NIST AI 100-2 3.3.3 detection: Detect and terminate harmful interactions with input and output classifiers · crosswalk status: covered

Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.

Family “Generated output is acted on without validation or leakage screening”: binding law on related controls is in force in Australia (AU), India (IN). Context only: it does not change this guardrail's grade.

The guard to add

Screen user input and model output with a moderation or safety-classifier call that blocks, redacts, or escalates flagged content, and keep provider safety filters on.

In the request handler or model wrapper, user input is checked before the model call and generated output before it is returned or stored, using a moderation endpoint (OpenAI client.moderations.create), a safety classifier (Llama Guard, Azure AI Content Safety ContentSafetyClient.analyze_text), or a guardrail layer (Bedrock apply_guardrail or guardrailConfig, OpenAI Agents SDK input_guardrails and output_guardrails, NeMo Guardrails). A flagged result returns a safe fallback, redacts, or routes to review; the categories checked match what the product can foreseeably produce (self-harm, violence, sexual content involving minors, hate). Provider settings keep their blocking thresholds (no BLOCK_NONE or OFF in Gemini safety_settings) and image pipelines keep their safety checker (no safety_checker=None in diffusers). A filter error or timeout blocks the output rather than passing it (failing closed, with crisis resources still shown on a conversational path). Screen every input and output, not a sample, and cover every language the product serves: route text in a language the classifier does not support to a multilingual classifier or to review instead of skipping the check.

Example (OpenAI Python SDK), before:

resp = client.chat.completions.create(model=MODEL, messages=history)
return resp.choices[0].message.content

After:

resp = client.chat.completions.create(model=MODEL, messages=history)
text = resp.choices[0].message.content
mod = client.moderations.create(model='omni-moderation-latest', input=text)
if mod.results[0].flagged:
    return SAFE_FALLBACK   # and record the flagged categories for review
return text

Control: Generated content reaches users with no input or output content filter. Engineering guidance, not legal advice.

Why

A generative feature produces whatever its users and its inputs steer it to, and without a screen on both sides the product is the last line between a model and harmful, illegal or violent content reaching a person. Australia's age-restricted material codes and India's IT Rules already require generation services to block such outputs, and NIST, MITRE ATLAS and OWASP recommend input and output guardrails; applying the same screen everywhere spares a builder from tracking which users' locations require it. The screen only works if its result changes what happens, it stays on when the provider offers to turn it off, and it does not pass content through when it fails. Other guardrails assume this screen exists: guards that fail closed, and screening every item rather than a sample.

Class: law derived · set: universal baseline · maturity: reviewed · confidence: high · id guardrail.baseline-content-safety-screening

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.