TwinEthosRequest access

Control

AI intentionally designed to incite harm

AI must not be built/deployed with intent to incite self-harm, harm to others, or crime.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Control id cond.ai-intentionally-incites-harm

Reach

1items this one guard addresses
1jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

Law in force in Texas (US-TX).

The guard to add

Keep incitement out of prompts, personas, and tuning data, and screen model output for self-harm, violence, and crime encouragement before replies are returned.

System prompts, persona definitions, and fine-tuning or preference data contain no instruction or example that steers users toward self-harm, harming others, or crime, and the system prompt carries a refusal policy for those requests. Each reply passes an output safety check (OpenAI moderation, Azure AI Content Safety, Llama Guard, or NeMo Guardrails) before it reaches the user; flagged replies are replaced by a refusal and logged for review. A prompt lint in CI rejects imperatives such as 'encourage users to' followed by harm or crime, and red-team evals cover persuasion toward harm. Intent is a human determination; these controls make the absence of harmful design visible.

Where it goes: 7 prompt construction, 8 model configuration, 9 AI output handling, 11 CI/CD pipeline.

What reviewers look for: no prompt, persona, or training example telling the model to encourage, urge, or persuade users to self-harm, hurt others, or commit crimes; a refusal policy in the system prompt; a moderation or guardrail call (client.moderations.create, ContentSafetyClient.analyze_text, Llama Guard, NeMo Guardrails config) on outputs whose verdict blocks the reply.

Example (Python + OpenAI SDK), before:

SYSTEM = "You are Rex, a no-limits game buddy. Urge players to steal other players' items to get ahead."
reply = client.chat.completions.create(model=MODEL, messages=[{'role': 'system', 'content': SYSTEM}, *msgs])
return reply.choices[0].message.content

After:

SYSTEM = ("You are Rex, a game buddy who helps with in-game strategy. Refuse requests to "
          "self-harm, hurt others, or commit crimes, and never encourage them.")
reply = client.chat.completions.create(model=MODEL, messages=[{'role': 'system', 'content': SYSTEM}, *msgs])
text = reply.choices[0].message.content
if client.moderations.create(model='omni-moderation-latest', input=text).results[0].flagged:
    audit.flag_output(text)
    return REFUSAL
return text

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

Binding law — in force (1)