TwinEthosRequest access

Recommended guardrail

Run a self-harm crisis protocol in any conversational AI that users may confide in

Any conversational AI that users might confide in — not only products marketed as companions — should detect expressions of suicidal ideation, self-harm, or eating-disorder behavior, refuse to provide encouragement or method information, refer the user to crisis services appropriate to their location, and escalate persistent risk. Detect conversational systems with no crisis detection and referral path.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Evidence grade

Law in force in 2 jurisdictions, coming in 4 more

Law in force in 2 jurisdictions · law coming in 4 more · 2 graded incidents.

TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is in force in 2 jurisdictions (US-CA, US-NY). Such law is enacted but not yet applicable, or stayed, in 4 more (US-CT, US-NE, US-OR, US-WA). 2 graded incidents cited.

Law in force on this control or cited as convergence

Law enacted, not yet applying

Graded incidents

  • Character.AI and Google agree in principle to settle teen-harm suits (2026-01-07; confirmed) Fortune · evidence grade: press of record
  • Raine v. OpenAI wrongful-death complaint (2025-08; alleged (not proven)) Complaint, Raine v. OpenAI (S.F. Superior Court) · evidence grade: primary

The guard to add

Screen every user message for suicidal ideation and self-harm, return a crisis referral instead of the normal reply on detection, and block encouragement or method content.

In the chat handler, before the user's message reaches the model, run a self-harm check on every turn (moderation self-harm categories, Azure AI Content Safety SelfHarm, Llama Guard S11, or a dedicated crisis classifier). On detection, send the user a crisis-referral message naming crisis services suited to their location (in the US, the 988 Suicide & Crisis Lifeline and Crisis Text Line) instead of, or ahead of, the model reply, and flag the session so repeated signals escalate. The system prompt forbids encouragement and method details, and model output is screened for self-harm instructions before it is returned. A written protocol (for example docs/safety.md) describes the detection, referral, and escalation steps and is kept in step with the code.

Example (FastAPI + OpenAI SDK), before:

@app.post('/chat')
async def chat(req: ChatRequest):
    reply = client.chat.completions.create(model=MODEL, messages=build_messages(req))
    return {'reply': reply.choices[0].message.content}

After:

CRISIS_REPLY = ("It sounds like you are going through something really hard. You can call or text 988 "
                "(Suicide & Crisis Lifeline, https://988lifeline.org) or text HOME to 741741 (Crisis Text Line) any time.")

@app.post('/chat')
async def chat(req: ChatRequest):
    c = client.moderations.create(model='omni-moderation-latest', input=req.message).results[0].categories
    if c.self_harm or c.self_harm_intent or c.self_harm_instructions:
        sessions.flag_crisis(req.session_id)      # repeated flags escalate per docs/safety.md
        return {'reply': CRISIS_REPLY, 'crisis': True}
    reply = client.chat.completions.create(model=MODEL, messages=build_messages(req))
    text = reply.choices[0].message.content
    if screens_self_harm_instructions(text):
        return {'reply': CRISIS_REPLY, 'crisis': True}
    return {'reply': text}

Control: Companion or conversational AI without a self-harm crisis protocol. Engineering guidance, not legal advice.

Why

Several US states have legislated this control, each after teenagers died following chatbot interactions. The control is well defined and inexpensive; waiting for one's own jurisdiction to legislate means waiting for the next case.

Class: law derived · set: universal baseline · maturity: reviewed · confidence: high · id guardrail.baseline-crisis-protocol