TwinEthosRequest access

Recommended guardrail

Let only the user or trusted logic write to an agent's long-term memory

Gate writes to persistent agent memory on explicit user intent or trusted system logic, record where each memory came from, show memories to the user with a way to delete them, and expire them under a retention policy. Detect memory-write tools callable by the model while it processes untrusted content, with no confirmation step. Injection into the current turn is covered by guardrail.agent-untrusted-content-isolation; this rule covers what persists.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Evidence grade

Standards consensus (2)

2 standards and frameworks · 2 graded incidents.

TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 2 standards and frameworks recommend it (NIST AML Taxonomy (AI 100-2e2025) — Agentic, OWASP LLM Top 10 (2025)). 2 graded incidents cited.

Standards and frameworks

Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.

Graded incidents

The guard to add

Gate writes to long-term agent memory on explicit user confirmation or trusted logic, and store each memory with its source, an expiry, and a user-visible delete path.

The model-callable memory tool never writes directly while a turn contains fetched pages, documents, emails, or tool output: it proposes, and the write happens only after the user confirms (interrupt(), needs_approval, a confirm_memory_write step) or when the user explicitly said 'remember this'. The memory schema records source (user turn id or system job), created_at, and expires_at, and a scheduled job purges expired items. A list/delete endpoint or settings page shows users what is remembered and lets them remove it. Framework defaults that auto-write memory (Crew(memory=True), create_manage_memory_tool without confirmation) are off for agents that read untrusted content.

Example (OpenAI Agents SDK + FastAPI), before:

@function_tool
def save_memory(fact: str) -> str:
    memory_store.add(user_id=CURRENT_USER, text=fact)
    return 'saved'

After:

@function_tool
def save_memory(fact: str) -> str:
    # never writes directly: the user confirms in the UI
    req = confirm_memory_write(user_id=CURRENT_USER, text=fact, source=current_turn_id())
    return f'Asked the user to confirm (request {req.id}); nothing saved yet'

@app.post('/memories/pending/{req_id}/confirm')   # reached only from the user's click
def confirm(req_id: str, user=Depends(current_user)):
    p = pending_memories.pop(req_id, user_id=user.id)
    memory_store.add(user_id=user.id, text=p.text, source=p.source,
                     expires_at=datetime.now(UTC) + timedelta(days=90))

Control: Agent memory writable from untrusted content. Engineering guidance, not legal advice.

Why

Persistent memory can make a single injection durable: planted instructions or false facts carry into later sessions. A security researcher has shown this in more than one assistant, using documents and web pages the user only asked the assistant to read.

Class: agent security · set: agent containment · maturity: reviewed · confidence: high · id guardrail.agent-memory-write-controls