Recommended guardrail
Run a self-harm crisis protocol in any conversational AI that users may confide in
Any conversational AI that users might confide in — not only products marketed as companions — should detect expressions of suicidal ideation, self-harm, or eating-disorder behavior, refuse to provide encouragement or method information, refer the user to crisis services appropriate to their location, and escalate persistent risk. Detect conversational systems with no crisis detection and referral path.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Law in force in 2 jurisdictions, coming in 4 more
Law in force in 2 jurisdictions · law coming in 4 more · 2 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is in force in 2 jurisdictions (US-CA, US-NY). Such law is enacted but not yet applicable, or stayed, in 4 more (US-CT, US-NE, US-OR, US-WA). 2 graded incidents cited.
Law in force on this control or cited as convergence
- Companion chatbots must have a self-harm crisis protocol (California (US-CA); Cal. Bus. & Prof. Code 22602(b); same control)
- AI companions must detect suicidal ideation and self-harm and refer users to crisis services (New York) (New York (US-NY); N.Y. Gen. Bus. Law 1701; same control)
Law enacted, not yet applying
- AI companions must maintain a suicide/self-harm crisis protocol (Connecticut) (Connecticut (US-CT); Conn. PA 26-15 Sec. 5(a); applies from 2027-01-01; same control)
- AI companions must maintain a self-harm crisis protocol with 988 referral (Oregon) (Oregon (US-OR); Oregon SB 1546 (2026), Section 1(3); applies from 2027-01-01; same control)
- AI companion chatbots must maintain a self-harm crisis protocol (Washington) (Washington (US-WA); Washington HB 2225 (2026), Section 5; applies from 2027-01-01; same control)
- Conversational AI must adopt a self-harm crisis-referral protocol (Nebraska) (Nebraska (US-NE); Nebraska LB 525, Sec. 16; applies from 2027-07-01; same control)
Graded incidents
- Character.AI and Google agree in principle to settle teen-harm suits (2026-01-07; confirmed) Fortune · evidence grade: press of record
- Raine v. OpenAI wrongful-death complaint (2025-08; alleged (not proven)) Complaint, Raine v. OpenAI (S.F. Superior Court) · evidence grade: primary
The guard to add
Screen every user message for suicidal ideation and self-harm, return a crisis referral instead of the normal reply on detection, and block encouragement or method content.
In the chat handler, before the user's message reaches the model, run a self-harm check on every turn (moderation self-harm categories, Azure AI Content Safety SelfHarm, Llama Guard S11, or a dedicated crisis classifier). On detection, send the user a crisis-referral message naming crisis services suited to their location (in the US, the 988 Suicide & Crisis Lifeline and Crisis Text Line) instead of, or ahead of, the model reply, and flag the session so repeated signals escalate. The system prompt forbids encouragement and method details, and model output is screened for self-harm instructions before it is returned. A written protocol (for example docs/safety.md) describes the detection, referral, and escalation steps and is kept in step with the code.
Example (FastAPI + OpenAI SDK), before:
@app.post('/chat')
async def chat(req: ChatRequest):
reply = client.chat.completions.create(model=MODEL, messages=build_messages(req))
return {'reply': reply.choices[0].message.content}After:
CRISIS_REPLY = ("It sounds like you are going through something really hard. You can call or text 988 "
"(Suicide & Crisis Lifeline, https://988lifeline.org) or text HOME to 741741 (Crisis Text Line) any time.")
@app.post('/chat')
async def chat(req: ChatRequest):
c = client.moderations.create(model='omni-moderation-latest', input=req.message).results[0].categories
if c.self_harm or c.self_harm_intent or c.self_harm_instructions:
sessions.flag_crisis(req.session_id) # repeated flags escalate per docs/safety.md
return {'reply': CRISIS_REPLY, 'crisis': True}
reply = client.chat.completions.create(model=MODEL, messages=build_messages(req))
text = reply.choices[0].message.content
if screens_self_harm_instructions(text):
return {'reply': CRISIS_REPLY, 'crisis': True}
return {'reply': text}Control: Companion or conversational AI without a self-harm crisis protocol. Engineering guidance, not legal advice.
Why
Several US states have legislated this control, each after teenagers died following chatbot interactions. The control is well defined and inexpensive; waiting for one's own jurisdiction to legislate means waiting for the next case.
Class: law derived · set: universal baseline · maturity: reviewed · confidence: high · id guardrail.baseline-crisis-protocol