Recommended guardrail
Screen model and agent output for personal and sensitive data before it leaves the trust boundary
Inspect generated output and agent-initiated transmissions for personal data, credentials, and confidential information before they reach a user outside the data's authorization scope or any external destination, and block or redact accordingly. Detect output paths — chat responses, summaries, emails, API calls, external posts — with no data-loss screening.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Recommended by 1 standard
1 standard or framework · 1 graded incident.
TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 1 standard or framework recommends it (NIST GenAI Profile (AI 600-1)). 1 graded incident cited.
Standards and frameworks
- GenAI outputs should be screened for PII/sensitive-data leakage (NIST GenAI Profile) (NIST GenAI Profile (AI 600-1); NIST AI 600-1 §2.4 (Data Privacy) + MP-4.1-009; same control)
Family “Generated output is acted on without validation or leakage screening”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
Graded incidents
- Slack AI indirect prompt injection (researcher disclosure) (2024-08; confirmed) PromptArmor (original researcher disclosure) · evidence grade: primary
The guard to add
Scan generated output for personal data, credentials, and secrets, and redact or block it before it is returned, posted, or sent beyond its authorized audience.
An output data-loss check in one shared helper that every outbound path uses: chat replies shown to users outside the data's scope, emails, Slack or webhook posts, tickets, and public pages. Run a PII and secret detector (Presidio AnalyzerEngine/AnonymizerEngine, llm_guard Sensitive, Azure AI Language PII, Google Cloud DLP) on the generated text, redact or block according to policy, and log the entity types found (not the values). Strip markdown images and links whose host is not on an allowlist, since an injected URL can carry data out.
Example (Python + Presidio + Slack SDK), before:
summary = completion.choices[0].message.content
slack_client.chat_postMessage(channel=SUPPORT_CHANNEL, text=summary)After:
from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
analyzer, anonymizer = AnalyzerEngine(), AnonymizerEngine()
def redact_pii(text: str) -> str:
findings = analyzer.analyze(text=text, language='en')
return anonymizer.anonymize(text=text, analyzer_results=findings).text
summary = redact_pii(completion.choices[0].message.content)
slack_client.chat_postMessage(channel=SUPPORT_CHANNEL, text=summary)Control: GenAI output path without PII/sensitive-data leakage detection. Engineering guidance, not legal advice.
Why
Assistants have been shown to be steerable into summarizing private material and routing it outward. Data-protection law makes the operator liable for the leak but does not require the screen that would prevent it.
Class: law derived · set: output integrity · maturity: reviewed · confidence: high · id guardrail.output-pii-leakage-screening