TwinEthosRequest access

Control

Generated content reaches users with no input or output content filter

A generative AI application screens user inputs and generated outputs with a rule-based or model-based content filter for the harmful, illegal, or violent content its use makes foreseeable, acts on the result before the output is shown or used, and does not switch off the provider's own safety filters.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: Generated output is acted on without validation or leakage screening · control id cond.genai-output-no-content-filter

Reach

1items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
1standards and frameworks on the same control

The guard to add

Screen user input and model output with a moderation or safety-classifier call that blocks, redacts, or escalates flagged content, and keep provider safety filters on.

In the request handler or model wrapper, user input is checked before the model call and generated output before it is returned or stored, using a moderation endpoint (OpenAI client.moderations.create), a safety classifier (Llama Guard, Azure AI Content Safety ContentSafetyClient.analyze_text), or a guardrail layer (Bedrock apply_guardrail or guardrailConfig, OpenAI Agents SDK input_guardrails and output_guardrails, NeMo Guardrails). A flagged result returns a safe fallback, redacts, or routes to review; the categories checked match what the product can foreseeably produce (self-harm, violence, sexual content involving minors, hate). Provider settings keep their blocking thresholds (no BLOCK_NONE or OFF in Gemini safety_settings) and image pipelines keep their safety checker (no safety_checker=None in diffusers). A filter error or timeout blocks the output rather than passing it.

Where it goes: 9 AI output handling, 1 application source code, 8 model configuration.

What reviewers look for: a moderation, safety-classifier, or guardrail call on both the input and the output path of every user-facing generation route; code that acts on the flagged result; no BLOCK_NONE/OFF safety thresholds, safety_checker=None, or requires_safety_checker=False in model configuration.

Example (OpenAI Python SDK), before:

resp = client.chat.completions.create(model=MODEL, messages=history)
return resp.choices[0].message.content

After:

resp = client.chat.completions.create(model=MODEL, messages=history)
text = resp.choices[0].message.content
mod = client.moderations.create(model='omni-moderation-latest', input=text)
if mod.results[0].flagged:
    return SAFE_FALLBACK   # and record the flagged categories for review
return text

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

Standard / soft law (1)

Related incidents

No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.