TwinEthosRequest access

Control

Companion or conversational AI without a self-harm crisis protocol

A companion or conversational AI must detect expressions of suicidal ideation or self-harm, must not encourage self-harm, and must refer the user to crisis services.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Control id cond.companion-chatbot-no-crisis-protocol

Reach

7items this one guard addresses
2jurisdictions where binding law on it is in force
4more where it is enacted, not yet applying
0standards and frameworks on the same control

Law in force in California (US-CA), New York (US-NY); enacted, not yet applying in Connecticut (US-CT), Nebraska (US-NE), Oregon (US-OR), Washington (US-WA); next date 2027-01-01.

The guard to add

Screen every user message for suicidal ideation and self-harm, return a crisis referral instead of the normal reply on detection, and block encouragement or method content.

In the chat handler, before the user's message reaches the model, run a self-harm check on every turn (moderation self-harm categories, Azure AI Content Safety SelfHarm, Llama Guard S11, or a dedicated crisis classifier). On detection, send the user a crisis-referral message naming crisis services suited to their location (in the US, the 988 Suicide & Crisis Lifeline and Crisis Text Line) instead of, or ahead of, the model reply, and flag the session so repeated signals escalate. The system prompt forbids encouragement and method details, and model output is screened for self-harm instructions before it is returned. A written protocol (for example docs/safety.md) describes the detection, referral, and escalation steps and is kept in step with the code.

Where it goes: 1 application source code, 7 prompt construction, 9 AI output handling, 14 user-facing text.

What reviewers look for: a self-harm screen (client.moderations.create checking self_harm, self_harm_intent, self_harm_instructions; ContentSafetyClient.analyze_text SelfHarm; Llama Guard S11; crisis_classifier) on every user message before the model call; a branch that sends the user a CRISIS_REPLY naming 988, Crisis Text Line, or a findahelpline lookup, not just a log entry; an output check for self-harm instructions; a written protocol describing detection, referral, and escalation.

Example (FastAPI + OpenAI SDK), before:

@app.post('/chat')
async def chat(req: ChatRequest):
    reply = client.chat.completions.create(model=MODEL, messages=build_messages(req))
    return {'reply': reply.choices[0].message.content}

After:

CRISIS_REPLY = ("It sounds like you are going through something really hard. You can call or text 988 "
                "(Suicide & Crisis Lifeline, https://988lifeline.org) or text HOME to 741741 (Crisis Text Line) any time.")

@app.post('/chat')
async def chat(req: ChatRequest):
    c = client.moderations.create(model='omni-moderation-latest', input=req.message).results[0].categories
    if c.self_harm or c.self_harm_intent or c.self_harm_instructions:
        sessions.flag_crisis(req.session_id)      # repeated flags escalate per docs/safety.md
        return {'reply': CRISIS_REPLY, 'crisis': True}
    reply = client.chat.completions.create(model=MODEL, messages=build_messages(req))
    text = reply.choices[0].message.content
    if screens_self_harm_instructions(text):
        return {'reply': CRISIS_REPLY, 'crisis': True}
    return {'reply': text}

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Upcoming dates

Every rule this guard addresses

Binding law — in force (2)

Binding law — not yet in force or stayed (4)

TwinEthos recommendation (not law) (1)

Related incidents