TwinEthos homeAPI access

Control

Content enters a retrieval index or agent memory without provenance, trust tier or sanitization

Content written to a retrieval index (vector store, search index) or to agent memory is normalized and stripped of invisible characters (zero-width, tag-block and variation-selector code points) and other hidden instructions, and carries its source, ingestion time and trust tier; content of different trust tiers is kept in separate indexes or collections, and retrieval and memory reads can filter and weight by trust tier.

Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.

Family: An AI agent's authority, reach, inputs, and components are not bounded and accountable · control id cond.indexed-or-remembered-content-without-provenance

Reach

1items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

Trust and provenance

How far the rules this guard addresses have been checked. Each rule links to its provision, with its citation, official text and its own panel.

This control
Audit-grade: meets all 3 checks of the TwinEthos audit standard that apply to it.
Lanes
TwinEthos recommendation (not law) 1
Data release
Data release 2026.10.03.4, data as of 3 Oct 2026, schema 0.3.10. This page also reflects corpus changes made after that release; they ship in the next one.
Legal review
None of the 1 rule has been reviewed by a lawyer; no TwinEthos rule has been legally reviewed yet. Treat each as research to check against the official text; it is not legal advice.
Audit standard
1 of 1 rule audit-grade. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
Detectors
1 detector, all experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify. Each provision lists its detectors' known limits.

The guard to add

Sanitize content and label its source and trust tier before it enters an index or memory; keep trust tiers apart.

In every ingest job and memory writer, before add_documents, from_documents, VectorStoreIndex.from_documents, collection.add, upsert, Memory.add or store.put: normalize text (unicodedata.normalize('NFKC', ...)), strip zero-width (U+200B to U+200D, U+2060), tag-block (U+E0000 to U+E007F) and variation-selector (U+FE00 to U+FE0F) characters, and attach metadata with the source, the ingestion time and a trust tier (operator, verified partner, user-supplied, public web). Write each trust tier to its own collection or namespace, never feed the agent's own outputs back into trusted memory, and let retrieval filter or down-weight by trust tier so untrusted text cannot outrank the operator's own.

Where it goes: 1 application source code, 7 prompt construction, 15 agent action surface.

What reviewers look for: a sanitize step and trust_tier (or equivalent) metadata on the path from a fetched page, upload or message to the index or memory write; separate collections or namespaces for untrusted sources; retrieval calls that filter on the trust tier.

Example (Python + LangChain + Chroma), before:

docs = WebBaseLoader(url).load()
vectorstore.add_documents(docs)

After:

INVISIBLE = re.compile('[\u200b-\u200d\u2060\ufe00-\ufe0f\U000e0000-\U000e007f]')

docs = WebBaseLoader(url).load()
for d in docs:
    d.page_content = INVISIBLE.sub('', unicodedata.normalize('NFKC', d.page_content))
    d.metadata.update(source=url, ingested_at=now_iso(), trust_tier='public_web')
public_web_store.add_documents(docs)   # its own collection, never the operator's

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

TwinEthos recommendation (not law) (1)

Related incidents

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.