Control
Content enters a retrieval index or agent memory without provenance, trust tier or sanitization
Content written to a retrieval index (vector store, search index) or to agent memory is normalized and stripped of invisible characters (zero-width, tag-block and variation-selector code points) and other hidden instructions, and carries its source, ingestion time and trust tier; content of different trust tiers is kept in separate indexes or collections, and retrieval and memory reads can filter and weight by trust tier.
Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.
Reach
Trust and provenance
How far the rules this guard addresses have been checked. Each rule links to its provision, with its citation, official text and its own panel.
- This control
- Audit-grade: meets all 3 checks of the TwinEthos audit standard that apply to it.
- Lanes
- TwinEthos recommendation (not law) 1
- Data release
- Data release 2026.10.03.4, data as of 3 Oct 2026, schema 0.3.10. This page also reflects corpus changes made after that release; they ship in the next one.
- Legal review
- None of the 1 rule has been reviewed by a lawyer; no TwinEthos rule has been legally reviewed yet. Treat each as research to check against the official text; it is not legal advice.
- Audit standard
- 1 of 1 rule audit-grade. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
- 1 detector, all experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify. Each provision lists its detectors' known limits.
The guard to add
Sanitize content and label its source and trust tier before it enters an index or memory; keep trust tiers apart.
In every ingest job and memory writer, before add_documents, from_documents, VectorStoreIndex.from_documents, collection.add, upsert, Memory.add or store.put: normalize text (unicodedata.normalize('NFKC', ...)), strip zero-width (U+200B to U+200D, U+2060), tag-block (U+E0000 to U+E007F) and variation-selector (U+FE00 to U+FE0F) characters, and attach metadata with the source, the ingestion time and a trust tier (operator, verified partner, user-supplied, public web). Write each trust tier to its own collection or namespace, never feed the agent's own outputs back into trusted memory, and let retrieval filter or down-weight by trust tier so untrusted text cannot outrank the operator's own.
Where it goes: 1 application source code, 7 prompt construction, 15 agent action surface.
What reviewers look for: a sanitize step and trust_tier (or equivalent) metadata on the path from a fetched page, upload or message to the index or memory write; separate collections or namespaces for untrusted sources; retrieval calls that filter on the trust tier.
Example (Python + LangChain + Chroma), before:
docs = WebBaseLoader(url).load()
vectorstore.add_documents(docs)After:
INVISIBLE = re.compile('[\u200b-\u200d\u2060\ufe00-\ufe0f\U000e0000-\U000e007f]')
docs = WebBaseLoader(url).load()
for d in docs:
d.page_content = INVISIBLE.sub('', unicodedata.normalize('NFKC', d.page_content))
d.metadata.update(source=url, ingested_at=now_iso(), trust_tier='public_web')
public_web_store.add_documents(docs) # its own collection, never the operator'sEngineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
TwinEthos recommendation (not law) (1)
- Everywhere (*)
- Sanitize and label the source and trust tier of everything written to an index or agent memory TwinEthos derivation — guardrail.agent-ingest-provenance-and-sanitization
Related incidents
- Gemini long-term memory poisoned through delayed tool invocation (2024-12; confirmed). Johann Rehberger showed that hidden instructions in a document Gemini summarized could plant a trigger so that, when the user later typed an ordinary word, Gemini saved false long-term memories as if the user had asked. He reported it to Google in December 2024 and published in February 2025; he reports Google assessed it as low likelihood and low impact. Source: Embrace The Red (Johann Rehberger, original researcher disclosure) · evidence grade: primary · cited by Sanitize and label the source and trust tier of everything written to an index or agent memory
- Slack AI indirect prompt injection (researcher disclosure) (2024-08; confirmed). Researchers showed that an instruction planted in a public Slack channel could make Slack AI leak private-channel data through a crafted link. Salesforce patched the issue and reported no evidence of unauthorized access to customer data. Source: PromptArmor (original researcher disclosure) · evidence grade: primary · cited by Sanitize and label the source and trust tier of everything written to an index or agent memory
- Prompt injection wrote persistent attacker instructions into ChatGPT memory (SpAIware) (2024-05; confirmed). Researcher Johann Rehberger showed that an untrusted website or document could use prompt injection to store attacker-chosen instructions in ChatGPT's persistent memory, which then carried into later chats and, in the macOS app, exfiltrated user data continuously. OpenAI fixed the exfiltration channel in September 2024; Rehberger says untrusted content can still write memories. Source: Embrace The Red (Johann Rehberger, original researcher disclosure) · evidence grade: primary · cited by Sanitize and label the source and trust tier of everything written to an index or agent memory
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.