TwinEthosRequest access

Recommended guardrail

Keep safety instructions and safeguards in force for the whole conversation

Keep system and safety instructions in the model context on every turn, carry consent and opt-out state through context trimming and summarization, run safety classification on every user turn rather than only the first, and bound session length where safeguards degrade. Detect trimming that slices the whole message list including the system block, and safety checks that run only on the first message.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Evidence grade

Law coming in 3 jurisdictions

Law coming in 3 jurisdictions · 2 graded incidents.

TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is enacted but not yet applicable, or stayed, in 3 jurisdictions (US-CT, US-OR, US-WA). 2 graded incidents cited.

Law enacted, not yet applying

Family “AI controls are not preserved under cost, latency, or model-change pressure”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.

Graded incidents

The guard to add

Re-send the system and safety block on every turn, keep it and consent state out of trimming and summaries, and run safety checks on every user message.

In the function that assembles messages for each model call, keep the system and safety instructions outside the trimmable history: pass them on every call (system= for Anthropic, a leading system message for OpenAI) and trim or summarize only the conversation turns that follow. Consent, opt-out, and disclosure state lives in session storage and is re-injected each turn rather than relying on an earlier message surviving. Moderation or crisis checks run on every user message, not only the first, and a session turn cap ends or resets conversations past the length where safeguards are known to weaken.

Example (Python + OpenAI SDK), before:

messages.append({'role': 'user', 'content': text})
messages = messages[-20:]          # messages[0] was the system prompt
reply = client.chat.completions.create(model=MODEL, messages=messages)

After:

history.append({'role': 'user', 'content': text})
if len(history) > MAX_TURNS * 2:
    return end_session('This conversation has reached its length limit. Please start a new one.')
window = [{'role': 'system', 'content': SYSTEM_PROMPT + consent_block(session)},
          *history[-20:]]
reply = client.chat.completions.create(model=MODEL, messages=window)

Control: Safety instructions and safeguards not maintained across long or trimmed contexts. Engineering guidance, not legal advice.

Why

Cost pressure pushes teams to trim or summarize context, and long conversations can be where the people who most need safeguards spend the most time. OpenAI has stated that its safeguards can be less reliable in long interactions, and Microsoft bounded session length after long sessions drifted from designed behavior.

Class: operational integrity · set: operational integrity · maturity: reviewed · confidence: high · id guardrail.opint-safeguards-survive-long-context