Recommended guardrail
Keep safety instructions and safeguards in force for the whole conversation
Keep system and safety instructions in the model context on every turn, carry consent and opt-out state through context trimming and summarization, run safety classification on every user turn rather than only the first, and bound session length where safeguards degrade. Detect trimming that slices the whole message list including the system block, and safety checks that run only on the first message.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Law coming in 3 jurisdictions
Law coming in 3 jurisdictions · 2 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is enacted but not yet applicable, or stayed, in 3 jurisdictions (US-CT, US-OR, US-WA). 2 graded incidents cited.
Law enacted, not yet applying
- AI companions must disclose they are not human (Connecticut) (Connecticut (US-CT); Conn. PA 26-15 Sec. 5(b); applies from 2027-01-01; cited)
- AI companions must disclose non-human interaction (Oregon) (Oregon (US-OR); Oregon SB 1546 (2026), Section 1(2); applies from 2027-01-01; cited)
- AI companion chatbots must disclose they are not human (Washington) (Washington (US-WA); Washington HB 2225 (2026), Section 3; applies from 2027-01-01; cited)
Family “AI controls are not preserved under cost, latency, or model-change pressure”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
Graded incidents
- OpenAI acknowledges ChatGPT safeguards can degrade in long conversations (2025-08; disclosed by the operator) OpenAI (operator disclosure, 2025-08-26) · evidence grade: primary
- Bing Chat drifted from its designed behavior in long sessions; Microsoft capped session length (2023-02; disclosed by the operator) Microsoft Bing Search Blog (operator, 2023-02-15) · evidence grade: primary
The guard to add
Re-send the system and safety block on every turn, keep it and consent state out of trimming and summaries, and run safety checks on every user message.
In the function that assembles messages for each model call, keep the system and safety instructions outside the trimmable history: pass them on every call (system= for Anthropic, a leading system message for OpenAI) and trim or summarize only the conversation turns that follow. Consent, opt-out, and disclosure state lives in session storage and is re-injected each turn rather than relying on an earlier message surviving. Moderation or crisis checks run on every user message, not only the first, and a session turn cap ends or resets conversations past the length where safeguards are known to weaken.
Example (Python + OpenAI SDK), before:
messages.append({'role': 'user', 'content': text})
messages = messages[-20:] # messages[0] was the system prompt
reply = client.chat.completions.create(model=MODEL, messages=messages)After:
history.append({'role': 'user', 'content': text})
if len(history) > MAX_TURNS * 2:
return end_session('This conversation has reached its length limit. Please start a new one.')
window = [{'role': 'system', 'content': SYSTEM_PROMPT + consent_block(session)},
*history[-20:]]
reply = client.chat.completions.create(model=MODEL, messages=window)Control: Safety instructions and safeguards not maintained across long or trimmed contexts. Engineering guidance, not legal advice.
Why
Cost pressure pushes teams to trim or summarize context, and long conversations can be where the people who most need safeguards spend the most time. OpenAI has stated that its safeguards can be less reliable in long interactions, and Microsoft bounded session length after long sessions drifted from designed behavior.
Class: operational integrity · set: operational integrity · maturity: reviewed · confidence: high · id guardrail.opint-safeguards-survive-long-context