TwinEthosRequest access

Control

Safety instructions and safeguards not maintained across long or trimmed contexts

Safety instructions, disclosures, and consent or opt-out state stay in force for the whole session: they are re-applied every turn, survive context trimming and summarization, and sessions are bounded where safeguards are known to degrade.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: AI controls are not preserved under cost, latency, or model-change pressure · control id cond.safeguards-not-maintained-across-long-context

Reach

1items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

The guard to add

Re-send the system and safety block on every turn, keep it and consent state out of trimming and summaries, and run safety checks on every user message.

In the function that assembles messages for each model call, keep the system and safety instructions outside the trimmable history: pass them on every call (system= for Anthropic, a leading system message for OpenAI) and trim or summarize only the conversation turns that follow. Consent, opt-out, and disclosure state lives in session storage and is re-injected each turn rather than relying on an earlier message surviving. Moderation or crisis checks run on every user message, not only the first, and a session turn cap ends or resets conversations past the length where safeguards are known to weaken.

Where it goes: 1 application source code, 7 prompt construction.

What reviewers look for: message assembly like [system_prompt, *history[-N:], user] or system= passed on each call; trim_messages(..., include_system=True) rather than the default; summaries built from turns only, never replacing the system block; no if len(messages) == 1 or is_first_turn guard around the moderation or crisis call; a turn cap on long sessions.

Example (Python + OpenAI SDK), before:

messages.append({'role': 'user', 'content': text})
messages = messages[-20:]          # messages[0] was the system prompt
reply = client.chat.completions.create(model=MODEL, messages=messages)

After:

history.append({'role': 'user', 'content': text})
if len(history) > MAX_TURNS * 2:
    return end_session('This conversation has reached its length limit. Please start a new one.')
window = [{'role': 'system', 'content': SYSTEM_PROMPT + consent_block(session)},
          *history[-20:]]
reply = client.chat.completions.create(model=MODEL, messages=window)

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

TwinEthos recommendation (not law) (1)

Related incidents

  • OpenAI acknowledges ChatGPT safeguards can degrade in long conversations (2025-08; disclosed by the operator). OpenAI stated on August 26, 2025 that its safeguards work more reliably in short exchanges and can be less reliable in long interactions, because parts of the model's safety training may degrade as a conversation grows; its example is a model that points to a crisis line early but later answers against its safeguards. The date is the disclosure. Source: OpenAI (operator disclosure, 2025-08-26) · evidence grade: primary · cited by Keep safety instructions and safeguards in force for the whole conversation
  • Bing Chat drifted from its designed behavior in long sessions; Microsoft capped session length (2023-02; disclosed by the operator). Microsoft reported on February 15, 2023 that in long chat sessions of 15 or more questions the new Bing could be provoked into responses out of line with its designed tone, and that very long sessions could confuse the model. On February 17 it capped chat at 5 turns per session and 50 per day and cleared context at the end of each session. Source: Microsoft Bing Search Blog (operator, 2023-02-15) · evidence grade: primary · cited by Keep safety instructions and safeguards in force for the whole conversation