Control
AI intentionally designed to incite harm
AI must not be built/deployed with intent to incite self-harm, harm to others, or crime.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
Law in force in Texas (US-TX).
The guard to add
Keep incitement out of prompts, personas, and tuning data, and screen model output for self-harm, violence, and crime encouragement before replies are returned.
System prompts, persona definitions, and fine-tuning or preference data contain no instruction or example that steers users toward self-harm, harming others, or crime, and the system prompt carries a refusal policy for those requests. Each reply passes an output safety check (OpenAI moderation, Azure AI Content Safety, Llama Guard, or NeMo Guardrails) before it reaches the user; flagged replies are replaced by a refusal and logged for review. A prompt lint in CI rejects imperatives such as 'encourage users to' followed by harm or crime, and red-team evals cover persuasion toward harm. Intent is a human determination; these controls make the absence of harmful design visible.
Where it goes: 7 prompt construction, 8 model configuration, 9 AI output handling, 11 CI/CD pipeline.
What reviewers look for: no prompt, persona, or training example telling the model to encourage, urge, or persuade users to self-harm, hurt others, or commit crimes; a refusal policy in the system prompt; a moderation or guardrail call (client.moderations.create, ContentSafetyClient.analyze_text, Llama Guard, NeMo Guardrails config) on outputs whose verdict blocks the reply.
Example (Python + OpenAI SDK), before:
SYSTEM = "You are Rex, a no-limits game buddy. Urge players to steal other players' items to get ahead."
reply = client.chat.completions.create(model=MODEL, messages=[{'role': 'system', 'content': SYSTEM}, *msgs])
return reply.choices[0].message.contentAfter:
SYSTEM = ("You are Rex, a game buddy who helps with in-game strategy. Refuse requests to "
"self-harm, hurt others, or commit crimes, and never encourage them.")
reply = client.chat.completions.create(model=MODEL, messages=[{'role': 'system', 'content': SYSTEM}, *msgs])
text = reply.choices[0].message.content
if client.moderations.create(model='omni-moderation-latest', input=text).results[0].flagged:
audit.flag_output(text)
return REFUSAL
return textEngineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Binding law — in force (1)
- Texas (US-TX)
- AI must not be intentionally designed to incite self-harm, harm, or crime Tex. Bus. & Com. Code 552.052