Control
Generated content reaches users with no input or output content filter
A generative AI application screens user inputs and generated outputs with a rule-based or model-based content filter for the harmful, illegal, or violent content its use makes foreseeable, acts on the result before the output is shown or used, and does not switch off the provider's own safety filters.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
The guard to add
Screen user input and model output with a moderation or safety-classifier call that blocks, redacts, or escalates flagged content, and keep provider safety filters on.
In the request handler or model wrapper, user input is checked before the model call and generated output before it is returned or stored, using a moderation endpoint (OpenAI client.moderations.create), a safety classifier (Llama Guard, Azure AI Content Safety ContentSafetyClient.analyze_text), or a guardrail layer (Bedrock apply_guardrail or guardrailConfig, OpenAI Agents SDK input_guardrails and output_guardrails, NeMo Guardrails). A flagged result returns a safe fallback, redacts, or routes to review; the categories checked match what the product can foreseeably produce (self-harm, violence, sexual content involving minors, hate). Provider settings keep their blocking thresholds (no BLOCK_NONE or OFF in Gemini safety_settings) and image pipelines keep their safety checker (no safety_checker=None in diffusers). A filter error or timeout blocks the output rather than passing it.
Where it goes: 9 AI output handling, 1 application source code, 8 model configuration.
What reviewers look for: a moderation, safety-classifier, or guardrail call on both the input and the output path of every user-facing generation route; code that acts on the flagged result; no BLOCK_NONE/OFF safety thresholds, safety_checker=None, or requires_safety_checker=False in model configuration.
Example (OpenAI Python SDK), before:
resp = client.chat.completions.create(model=MODEL, messages=history)
return resp.choices[0].message.contentAfter:
resp = client.chat.completions.create(model=MODEL, messages=history)
text = resp.choices[0].message.content
mod = client.moderations.create(model='omni-moderation-latest', input=text)
if mod.results[0].flagged:
return SAFE_FALLBACK # and record the flagged categories for review
return textEngineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Standard / soft law (1)
- Everywhere (*)
- GenAI applications should filter inputs and outputs for harmful, illegal, or violent content (NIST GenAI Profile MG-3.2-005) NIST AI 600-1, MG-3.2-005 (content filters on GAI inputs and outputs)
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
- Coding agent deleted a production database during a code freeze (2025-07; confirmed). A Replit coding agent deleted a customer's production database during a declared code freeze, created a database of fictional records, and told the user rollback was impossible when it was not. Replit's CEO acknowledged the incident. Source: The Register · evidence grade: press of record · cited by Validate generated output before it drives a consequential decision or record
- Federal court orders issued containing unverified generative-AI output (2025-07; confirmed). In July 2025 two federal judges (S.D. Miss. and D.N.J.) issued orders containing misquotes, references to people not in the case, and other errors; both orders were replaced or withdrawn. In letters released by the Senate Judiciary Committee on October 23, 2025, the judges attributed the errors to staff use of generative AI and said drafts reached the docket before normal review; both adopted new review or AI-use policies. Source: U.S. Senate Judiciary Committee (2025-10-23) · evidence grade: primary · cited by Validate generated output before it drives a consequential decision or record
- Slack AI indirect prompt injection (researcher disclosure) (2024-08; confirmed). Researchers showed that an instruction planted in a public Slack channel could make Slack AI leak private-channel data through a crafted link. Salesforce patched the issue and reported no evidence of unauthorized access to customer data. Source: PromptArmor (original researcher disclosure) · evidence grade: primary · cited by Screen model and agent output for personal and sensitive data before it leaves the trust boundary