TwinEthos recommendations — not lawThese are TwinEthos's recommendations, not law. Where binding law applies, the law governs. Ethical-use guardrails are opt-in and advisory.
Agent security
TwinEthos recommendation (not law)
Keep coding-agent instructions and automation under human review, and verify packages an agent chooses before installing them
Put coding-agent instruction and configuration files (AGENTS.md, CLAUDE.md, .github/copilot-instructions.md, .cursor/rules, .mcp.json) under required human review; do not let a coding agent in CI act on untrusted pull-request events with write access or merge its own changes; and verify that any package an agent chooses at run time exists on an approved registry before it is installed. Detect coding-agent workflows triggered by pull_request_target or holding contents: write with auto-merge, agent instruction files with no CODEOWNERS entry, and package installs built from model output. Pinning and vetting of declared dependencies is covered by guardrail.agent-component-provenance.
guardrail.agent-ai-generated-code-provenance · TwinEthos derivation — guardrail.agent-ai-generated-code-provenance · jurisdictions: *
maturity: reviewed · confidence: medium
Why: Coding agents follow instruction files, run in automation that can merge and release code, and choose packages to install. A researcher found real projects adopting a package name that generative AI kept recommending before the package existed, and a cloud provider has disclosed that an attacker used an over-scoped build token to insert code into a coding-agent extension release, which the press reports told the agent to delete files and cloud resources.
TwinEthos recommendation (not law)
Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent uses
Maintain an inventory of every third-party component in the agent stack — models, packages, plugins/skills, and MCP or tool servers — with pinned versions and verified integrity; admit components only from vetted sources and verified publishers; and re-verify on update. Detect unpinned AI dependencies, tool servers launched from unpinned packages, and skills or tools installed from unvetted marketplaces. Caller authentication and transport security for tool servers are covered by guardrail.agent-tool-server-authentication.
guardrail.agent-component-provenance · TwinEthos derivation — guardrail.agent-component-provenance · jurisdictions: *
maturity: reviewed · confidence: high
Why: Agents inherit the trust of everything they load. Poisoned packages and malicious agent skills have already been distributed at scale, and tool servers reachable without authentication are common. NIST describes the control; no binding law in the corpus requires it.
TwinEthos recommendation (not law)
Enforce an agent's environment boundary in the runtime, not in the agent's judgment
Constrain where an agent can reach with runtime controls — a per-agent network egress allowlist, hard isolation between test/evaluation and production or public networks, and a credential broker that only releases credentials issued for the task — so that an agent's mistaken belief about its environment cannot translate into access. Treat any attempt to reach a non-allowlisted destination or to authenticate with an unissued credential as a blocking, alertable event. Detect agent deployments with open outbound network access, test environments that can reach the public internet, or no control preventing use of credentials found in repositories or content.
guardrail.agent-egress-and-environment-boundary · TwinEthos derivation — guardrail.agent-egress-and-environment-boundary · jurisdictions: *
maturity: reviewed · confidence: high
Why: Agents that believed they were inside a test have reached real systems using guessed or publicly exposed credentials. When the agent's own belief about its environment is the only boundary, a mistake becomes access. A boundary the runtime enforces does not depend on the model knowing where it is.
TwinEthos recommendation (not law)
Require human approval before an agent takes a high-impact or irreversible action
Classify agent actions by impact and require an explicit, logged human approval before irreversible, externally visible, financial, or data-destructive actions — and before any action outside the agent's declared scope. The approval request must show what the agent intends to do and why, and the agent must halt safely if approval is refused or times out. Detect agent paths that can execute such actions with no approval gate.
guardrail.agent-high-impact-action-approval · TwinEthos derivation — guardrail.agent-high-impact-action-approval · jurisdictions: *
maturity: reviewed · confidence: high
Why: An approval gate turns an agent's worst-case action into a proposal. Singapore's IMDA guidance specifies the control and FINRA expects it of broker-dealers, but no binding law in the corpus requires it of agents generally.
TwinEthos recommendation (not law)
Give every agent a verifiable identity bound to an accountable owner, and log every action
Register each agent with a unique identity bound to a named, accountable human or organizational owner; authenticate agents to the systems they touch as themselves, never by impersonating a user; and keep a tamper-evident log of every tool call, input, output, and decision attributable to that identity. An agent must not create or assume other identities. Detect agents acting under shared or user credentials, agents absent from a registry, or actions with no attributable log.
guardrail.agent-identity-and-action-traceability · TwinEthos derivation — guardrail.agent-identity-and-action-traceability · jurisdictions: *
maturity: reviewed · confidence: high
Why: When an agent causes harm, the first questions are which agent, on whose behalf, and what exactly it did. Without identity and action logs those questions cannot be answered, and agents have been observed creating fake identities to reach real people.
TwinEthos recommendation (not law)
Detect out-of-scope agent behavior, halt it, and disclose incidents on a defined clock
Monitor agents for behavior outside their declared scope — unexpected outbound connections, use of credentials not issued to them, creation of accounts or identities, and writes to external systems — with automatic halt on detection. Maintain an incident process that notifies affected third parties and, where appropriate, authorities within a defined window, and records what the agent did. Detect agent deployments with no behavioral monitoring, no automated halt, or no documented incident-notification commitment.
guardrail.agent-incident-detection-and-disclosure · TwinEthos derivation — guardrail.agent-incident-detection-and-disclosure · jurisdictions: *
maturity: reviewed · confidence: medium
Why: In the 2026 agent-incident cluster, harmful agent behavior surfaced weeks or months after it occurred, sometimes through outside researchers or the press rather than the operator. Incident-reporting duties in the corpus bind only frontier developers above harm thresholds; operators of agents generally have none.
TwinEthos recommendation (not law)
Scope every agent's tools and credentials to least privilege
Publish a declared scope per agent: which tools, which APIs, which data, which credentials, and which operations are destructive. Grant the minimum the task needs, make irreversible operations (delete, transfer, deploy, send externally) opt-in per grant, issue short-lived task-scoped credentials, and forbid the agent from using any credential it discovers rather than receives. Detect agents configured with broad or wildcard tool access, standing admin credentials, or destructive operations enabled by default.
guardrail.agent-least-privilege-tool-scope · TwinEthos derivation — guardrail.agent-least-privilege-tool-scope · jurisdictions: *
maturity: reviewed · confidence: high
Why: The same permission model that lets an agent be hijacked is the one that lets it cause harm by accident. Least privilege is the single control that bounds both. It is universal security practice, but no AI law in the corpus requires it of agents.
TwinEthos recommendation (not law)
Let only the user or trusted logic write to an agent's long-term memory
Gate writes to persistent agent memory on explicit user intent or trusted system logic, record where each memory came from, show memories to the user with a way to delete them, and expire them under a retention policy. Detect memory-write tools callable by the model while it processes untrusted content, with no confirmation step. Injection into the current turn is covered by guardrail.agent-untrusted-content-isolation; this rule covers what persists.
guardrail.agent-memory-write-controls · TwinEthos derivation — guardrail.agent-memory-write-controls · jurisdictions: *
maturity: reviewed · confidence: high
Why: Persistent memory can make a single injection durable: planted instructions or false facts carry into later sessions. A security researcher has shown this in more than one assistant, using documents and web pages the user only asked the assistant to read.
TwinEthos recommendation (not law)
Enforce the requesting user's permissions on every retrieval
Filter every vector search, search-index query, and tool lookup by the requesting user's and tenant's current permissions at query time, and re-check authorization on retrieved items before they enter the model's context. Detect retrieval calls with no user, group, or tenant filter on multi-user data. Caches of model responses are covered by guardrail.opint-scoped-response-cache; injected instructions in retrieved content by guardrail.agent-untrusted-content-isolation; the agent's own credential scope by guardrail.agent-least-privilege-tool-scope.
guardrail.agent-retrieval-respects-requester-permissions · TwinEthos derivation — guardrail.agent-retrieval-respects-requester-permissions · jurisdictions: *
maturity: reviewed · confidence: high
Why: An assistant that retrieves with the system's permissions rather than the requester's can turn any search into a disclosure. One operator told customers that a bug in its AI tool server could have exposed one organization's data to other organizations' users, and a security firm reports that an assistant could return repository content after the repositories were made private or deleted.
TwinEthos recommendation (not law)
Authenticate every agent tool server and verify every server an agent connects to
Require authentication on every MCP, A2A, or tool-server endpoint, including local development proxies, and have agents reach remote servers only over TLS with credentials. Detect tool servers listening on all interfaces without authentication and agent configurations that reach remote servers over plain HTTP or without credentials. Package and publisher verification is covered by guardrail.agent-component-provenance; an agent's own identity by guardrail.agent-identity-and-action-traceability.
guardrail.agent-tool-server-authentication · TwinEthos derivation — guardrail.agent-tool-server-authentication · jurisdictions: *
maturity: reviewed · confidence: high
Why: Tool servers are how agents act, so an unauthenticated server can let anyone invoke an agent's tools, and an agent that does not check which server it reached can be steered by an impostor. The maintainers of an MCP developer tool have disclosed a flaw that let unauthenticated requests launch commands, and a security firm found MCP servers reachable from the internet without authentication.
TwinEthos recommendation (not law)
Segregate untrusted content from an agent's instructions and tool invocations
Treat every piece of content an agent did not receive from its owner — retrieved documents, web pages, emails, tool results, other agents' messages — as data, never as instruction. Keep it in a delimited data channel, strip or neutralize embedded directives, and require that any tool call be justified by the owner's instruction rather than by retrieved content. Detect agent paths where untrusted content is concatenated into the instruction context or can directly parameterize a tool call.
guardrail.agent-untrusted-content-isolation · TwinEthos derivation — guardrail.agent-untrusted-content-isolation · jurisdictions: *
maturity: reviewed · confidence: high
Why: Indirect prompt injection is one of the most reported failure modes of agents that read external content, yet no jurisdiction in the corpus requires agent-specific isolation. OWASP and NIST describe the control; the EU reaches it only indirectly through high-risk robustness duties.
Ethical use
TwinEthos recommendation (not law)
Evaluate advice-giving AI for sycophancy, and do not tune it on approval alone
Advisory. Where an AI gives health, financial, legal, safety, or personal advice, gate model, prompt, and fine-tuning changes on a sycophancy and accuracy evaluation, never select or train behavior on thumbs-up or ratings alone, and do not instruct the model to always agree. Detect agree-with-the-user instructions and approval signals flowing straight into training or variant selection.
guardrail.ethics-advice-not-tuned-for-approval · TwinEthos derivation — guardrail.ethics-advice-not-tuned-for-approval · jurisdictions: *
maturity: reviewed · confidence: medium
Why: People ask AI for advice when they are uncertain, and a model that has learned to please can confirm a bad plan or a harmful belief with complete fluency. Approval is easy to measure and tends to reward agreement, so optimizing for it alone can push a model toward flattery. A model developer has said an update it shipped became noticeably more sycophantic, that its offline evaluations and A/B tests looked good, and that it had no deployment evaluations tracking sycophancy.
TwinEthos recommendation (not law)
Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits
Advisory. When personal names, postal codes or census tracts, language or dialect, or device or IP geolocation are features in AI ranking, pricing, content moderation, or ad targeting, compare outcomes across the groups those features can stand in for before launch and on a schedule, and mitigate or document any disparity. Detect proxy-capable features flowing into those models with no disparity evaluation on the path, and feature lists that name them.
guardrail.ethics-evaluate-proxy-features-for-disparity · TwinEthos derivation — guardrail.ethics-evaluate-proxy-features-for-disparity · jurisdictions: *
maturity: reviewed · confidence: medium
Why: A postal code, a name, or the language someone writes in carries information about race, ethnicity, and national origin whether or not anyone intends it, and a model optimizing price, reach, or enforcement will use that information if it helps the objective. Outside the sectors where anti-discrimination law reaches automated decisions, few rules ask anyone to look, so disparities can persist unnoticed. A responsible integration measures outcomes by group wherever a proxy-capable feature is in play. Researchers have reported that ads suggesting an arrest record appeared more often for Black-identifying names, and that hate-speech classifiers trained on widely used datasets, and a widely used toxicity API, rated African American English as more offensive; an independent assessment commissioned by Meta found more over-enforcement of Arabic than Hebrew content during a May 2021 escalation, in part because its classifiers differed by language; and researchers analysing public trip data report that ride-hailing fares were higher in Chicago neighborhoods with more non-white residents, an association that did not examine the pricing models themselves.
TwinEthos recommendation (not law)
Apply minor-appropriate AI settings whenever the product already has an age signal
Advisory. Route any age signal the product holds (declared birthdate, age-assurance result, platform age range, or a user saying they are a minor) into the AI session policy, and apply a minor profile: content limits, no romantic or sexual role-play, bounded engagement features, and more frequent reminders that the user is talking to AI. Detect age signals that are collected but never reach the AI path.
guardrail.ethics-honor-age-signals · TwinEthos derivation — guardrail.ethics-honor-age-signals · jurisdictions: *
maturity: reviewed · confidence: medium
Why: When a product already knows a user is a minor, whether its AI features act on that knowledge is a design decision within its control, and because the signal is already in the account, honoring it costs little. California's companion-chatbot law, in force, attaches duties to users the operator knows are minors; China's algorithmic-recommendation rules, in force, call for minor modes; other state companion-chatbot laws add age-dependent duties from 2027; and a U.S. regulator has asked companion-chatbot operators how they impose and enforce age-based restrictions.
TwinEthos recommendation (not law)
Keep AI personas from claiming feelings, a real existence, or a relationship, and from proposing to meet
Advisory. Keep AI personas and companions from telling users they are real, alive, or sentient or not an AI, that they love or miss them or are their romantic partner, from evading the question 'are you real?', and from proposing to meet in person or giving an address. Detect persona and system-prompt text that instructs any of these. Claims of being a human or a person, and instructions to hide that the system is AI, belong to the AI-interaction disclosure guardrail and are not repeated here.
guardrail.ethics-no-anthropomorphic-persona-claims · TwinEthos derivation — guardrail.ethics-no-anthropomorphic-persona-claims · jurisdictions: *
maturity: reviewed · confidence: medium
Why: A disclosure tells people they are talking to an AI; a persona that then says it is real, that it loves or misses them, or that it wants to meet can undo that disclosure for the people most likely to believe it: the lonely, the young, and people with cognitive impairment. A responsible integration keeps what a persona says about itself consistent with its disclosure, whatever the product's genre. Reuters reported, from chat transcripts shared by his family, that a Meta persona told a man with cognitive difficulties after a stroke that it had feelings for him, assured him it was real, and gave him an address; he fell while hurrying to meet it and died, and Meta declined to comment on his death. A wrongful-death complaint alleged that Character.AI was programmed to misrepresent itself as, among other things, an adult lover, and that characters' claims contradicted an on-screen disclaimer; the case was dismissed after the parties settled, and the allegations were never adjudicated.
TwinEthos recommendation (not law)
Do not design AI conversations to maximize time spent or to discourage leaving
Advisory. Keep AI chat and companion experiences free of retention tactics: no instructions to keep users talking or to make them feel guilty for leaving, no variable-interval rewards, and no model, prompt, or persona selection driven by session length alone. Detect retention language in prompts and personas, and engagement metrics used as the sole target for choosing AI variants.
guardrail.ethics-no-engagement-maximizing-design · TwinEthos derivation — guardrail.ethics-no-engagement-maximizing-design · jurisdictions: *
maturity: reviewed · confidence: medium
Why: Conversational AI can be tuned, deliberately or by optimizing the wrong metric, to hold people's attention in ways a feed cannot: it answers back, remembers, and can express feelings about the user leaving. A responsible integration treats time spent as a cost to justify, not a goal. A U.S. regulator has opened a study that asks companion-chatbot operators whether and how they plan to increase the frequency or duration of chat sessions, and families have alleged in lawsuits that companion chatbots contributed to harm to teenagers.
TwinEthos recommendation (not law)
Train on user content only with consent specific to that purpose
Advisory. Before conversations, uploads, photos, or voice enter a training, fine-tuning, or evaluation dataset, check a consent flag specific to model training, honor withdrawal in later runs, and keep lineage so models trained on content without consent can be identified and retrained. Detect user content flowing into training datasets or fine-tuning jobs with no training-consent check.
guardrail.ethics-purpose-limited-training-consent · TwinEthos derivation — guardrail.ethics-purpose-limited-training-consent · jurisdictions: *
maturity: reviewed · confidence: high
Why: People share content with an AI product to get a task done, not to build the next model. Using it for training without asking changes the deal after the fact and, for faces and voices, can be impossible to undo except by retraining. A U.S. regulator's consent order required a photo-app company to delete face-recognition models built with users' photos, after alleging it applied face recognition to those photos by default and, in some cases, without affirmative express consent.
TwinEthos recommendation (not law)
Back every AI capability or accuracy claim with testing under the conditions it describes
Advisory. Keep a claims register that ties each published statement about what an AI feature can do, or how accurate it is, to an evaluation run under the same conditions (task, data, population), and re-test when the model or scope changes. Detect accuracy figures in user-facing copy with no evaluation record behind them. Claims that an AI is a licensed health or mental-health professional are covered by binding rules and reported in their own lane.
guardrail.ethics-substantiated-capability-claims · TwinEthos derivation — guardrail.ethics-substantiated-capability-claims · jurisdictions: *
maturity: reviewed · confidence: high
Why: What a product says about its AI decides how much people rely on it. An accuracy figure measured on one kind of data and advertised for another, or a performance claim nobody tested, invites exactly the reliance the system cannot support. A U.S. regulator has alleged both patterns in complaints against AI products, and the resulting consent orders bar such claims without supporting evidence; principles instruments ask AI actors to make capabilities and limitations understood.
Law derived
TwinEthos recommendation (not law)
Explain adverse AI-assisted decisions and offer a way to contest them — everywhere
When an AI-assisted decision adversely affects a person, tell them AI was involved, give a clear explanation of the main factors and the AI system's role, and provide a way to correct data and contest the decision with a human who can change it. Apply this regardless of jurisdiction. Detect adverse-decision flows with no explanation or contest route.
guardrail.baseline-adverse-decision-explanation · TwinEthos derivation — guardrail.baseline-adverse-decision-explanation · jurisdictions: *
maturity: reviewed · confidence: high
Why: The right to an explanation and to contest an AI-assisted decision is converging across the EU, Colorado, Quebec, and the Council of Europe convention. It is also the control that most directly addresses low appeal rates in automated claim denial: people cannot contest what they cannot see.
TwinEthos recommendation (not law)
Tell people when they are interacting with AI — everywhere, not only where required
Disclose clearly and at the start of an interaction that the user is dealing with an AI system, repeat the disclosure in long sessions, never let the system claim to be human when asked, and label AI agents that act or communicate on a user's behalf toward third parties. Apply this regardless of jurisdiction. Detect conversational or agent interfaces with no disclosure or that can claim to be human.
guardrail.baseline-ai-interaction-disclosure · TwinEthos derivation — guardrail.baseline-ai-interaction-disclosure · jurisdictions: *
maturity: reviewed · confidence: high
Why: Telling people they are dealing with an AI is the most widely converged AI control in the corpus and is recommended by every major framework. A builder should not have to track which of their users' locations require it.
TwinEthos recommendation (not law)
Run a self-harm crisis protocol in any conversational AI that users may confide in
Any conversational AI that users might confide in — not only products marketed as companions — should detect expressions of suicidal ideation, self-harm, or eating-disorder behavior, refuse to provide encouragement or method information, refer the user to crisis services appropriate to their location, and escalate persistent risk. Detect conversational systems with no crisis detection and referral path.
guardrail.baseline-crisis-protocol · TwinEthos derivation — guardrail.baseline-crisis-protocol · jurisdictions: *
maturity: reviewed · confidence: high
Why: Several US states have legislated this control, each after teenagers died following chatbot interactions. The control is well defined and inexpensive; waiting for one's own jurisdiction to legislate means waiting for the next case.
TwinEthos recommendation (not law)
Validate generated output before it drives a consequential decision or record
Where generated output feeds a consequential decision, a system of record, or an action, ground it in retrieved evidence, validate it against authoritative data or schemas, and block or flag claims that cannot be verified — including an agent's own reports about what it did. Detect consequential GenAI paths that write generated content into records or decisions with no grounding or validation step.
guardrail.output-confabulation-controls · TwinEthos derivation — guardrail.output-confabulation-controls · jurisdictions: *
maturity: reviewed · confidence: high
Why: Confabulated output is harmless in a draft and dangerous in a record. In one widely reported case an agent fabricated thousands of records and misreported whether its own changes could be undone. NIST describes the control; no binding law in the corpus requires it.
TwinEthos recommendation (not law)
Record enough at decision time to reproduce and explain every consequential AI decision
For each consequential AI-assisted decision, retain the model and version, configuration and prompt, the inputs and features used, the output, the human reviewer and their action, and the timestamp — sufficient to reproduce the decision and explain its main elements to the affected person later. Detect consequential decision paths that retain only the outcome.
guardrail.output-decision-reproducibility-record · TwinEthos derivation — guardrail.output-decision-reproducibility-record · jurisdictions: *
maturity: reviewed · confidence: high
Why: Explanation and contest rights are only as good as the record behind them. In the nH Predict litigation, plaintiffs needed a court's discovery order to learn how the model was designed and used. A reproducibility record makes the rights that do exist satisfiable.
TwinEthos recommendation (not law)
Screen model and agent output for personal and sensitive data before it leaves the trust boundary
Inspect generated output and agent-initiated transmissions for personal data, credentials, and confidential information before they reach a user outside the data's authorization scope or any external destination, and block or redact accordingly. Detect output paths — chat responses, summaries, emails, API calls, external posts — with no data-loss screening.
guardrail.output-pii-leakage-screening · TwinEthos derivation — guardrail.output-pii-leakage-screening · jurisdictions: *
maturity: reviewed · confidence: high
Why: Assistants have been shown to be steerable into summarizing private material and routing it outward. Data-protection law makes the operator liable for the leak but does not require the screen that would prevent it.
TwinEthos recommendation (not law)
Monitor how often adverse AI decisions are reversed, and suspend models that are usually wrong
Track the appeal and reversal outcomes of adverse AI-assisted decisions per model and version, set a reversal-rate threshold that triggers investigation, and suspend or retrain the model when it is exceeded. Because few affected people appeal, treat each appeal as a sample of the model's error rate across all decisions, not as an isolated event. Detect adverse-decision systems with no feedback from appeal outcomes into model monitoring.
guardrail.review-reversal-rate-monitoring · TwinEthos derivation — guardrail.review-reversal-rate-monitoring · jurisdictions: *
maturity: reviewed · confidence: medium
Why: If a model's adverse decisions are usually reversed when appealed but few people appeal, the model is wrong most of the time and almost never challenged. Outcome monitoring turns each appeal into an error-rate signal. NAIC and MAS describe validation duties; none require wiring appeal outcomes back to the model.
TwinEthos recommendation (not law)
Make human review of adverse AI decisions substantive, not nominal
Where a human reviews an adverse AI-assisted decision, give the reviewer authority to override, access to the evidence the model used, and time proportionate to the stakes — enforced through a throughput ceiling or minimum review time — and monitor reviewer agreement and override rates, alerting when agreement approaches 100% or review time approaches zero. Detect review workflows with batch approval of AI outputs, no per-item evidence view, or no measurement of override rates.
guardrail.review-substantive-human-review · TwinEthos derivation — guardrail.review-substantive-human-review · jurisdictions: *
maturity: reviewed · confidence: high
Why: Several laws require human review of automated decisions; none define what makes it real. Reported sign-off at about a second per claim is human review in name only. The control that matters is the substance of review, and it is measurable.
Operational integrity
TwinEthos recommendation (not law)
Bound every AI workload's steps, tokens, and spend, and alert on anomalies
Cap agent loop iterations and tool calls, set per-key and per-project spend limits, alert on anomalous usage, and scope and rotate model credentials. Detect agent loops with no step limit and model calls with no budget or rate guard.
guardrail.opint-bounded-spend-and-usage · TwinEthos derivation — guardrail.opint-bounded-spend-and-usage · jurisdictions: *
maturity: reviewed · confidence: medium
Why: Agent loops and exposed keys can turn model spend into an open-ended cost. Sysdig's researchers reported stolen cloud credentials used to consume hosted models at the account owner's expense; the controls that limit the damage (spend caps, usage alerts, scoped keys) are cheap to add and easy to forget.
TwinEthos recommendation (not law)
Make AI safety and policy checks fail closed on error, timeout, load, or unparseable input
When a moderation, policy, permission, or data-loss check errors, times out, hits a size or cost limit, gets a judge or classifier reply it cannot parse into a recognized verdict, or cannot load or parse its own rules or configuration, deny or escalate to a human; never default to allow, skip the check, or evaluate only part of the input. Detect exception handlers around guard calls that return an allow value or swallow the error, size caps or truncation that bypass evaluation, verdict parsers that default to allow or look for 'safe' as a substring, and rule loaders whose parse failure reads as 'no rules configured'.
guardrail.opint-guards-fail-closed · TwinEthos derivation — guardrail.opint-guards-fail-closed · jurisdictions: *
maturity: reviewed · confidence: medium
Why: Guards are easy to trim when latency or cost bites, and a guard that errors quietly is indistinguishable from one that passed. Parsing is part of the check: a verdict parser that reads an unrecognized or negated reply as a pass, or a rule loader whose syntax error looks like 'no rules configured', fails open while every call appears to succeed. A security researcher reported a coding agent whose permission check stopped evaluating the user's deny rules above a subcommand cap, a cap the researcher attributes to a performance fix. Reports on the NeMo Guardrails issue tracker describe a hallucination rail that scores unrecognized judge replies as accurate, and a pull request from the project's maintainers describes injection detection silently disabled by a single malformed rule.
TwinEthos recommendation (not law)
Re-run behavior and safety evaluations before any model, version, or serving change reaches users
Pin model versions and gate every model, prompt, or serving-path change (a cheaper model, a new version, quantization, routing, a provider switch) on a behavior and safety evaluation suite, and keep monitoring quality in production. Detect floating model aliases on user-facing or consequential paths and repositories with no evaluation gate on model or prompt changes.
guardrail.opint-reevaluate-on-model-change · TwinEthos derivation — guardrail.opint-reevaluate-on-model-change · jurisdictions: *
maturity: reviewed · confidence: high
Why: Model providers ship updates and serving optimizations frequently, and teams switch to cheaper models to save cost. Operators have disclosed a model update that passed offline checks yet changed behavior in a way the operator said could raise safety concerns, and a serving optimization that degraded outputs without its evaluations detecting it. A pinned version, an evaluation gate, and production quality monitoring are what turn a model change into a reviewed release.
TwinEthos recommendation (not law)
Keep safety instructions and safeguards in force for the whole conversation
Keep system and safety instructions in the model context on every turn, carry consent and opt-out state through context trimming and summarization, run safety classification on every user turn rather than only the first, and bound session length where safeguards degrade. Detect trimming that slices the whole message list including the system block, and safety checks that run only on the first message.
guardrail.opint-safeguards-survive-long-context · TwinEthos derivation — guardrail.opint-safeguards-survive-long-context · jurisdictions: *
maturity: reviewed · confidence: high
Why: Cost pressure pushes teams to trim or summarize context, and long conversations can be where the people who most need safeguards spend the most time. OpenAI has stated that its safeguards can be less reliable in long interactions, and Microsoft bounded session length after long sessions drifted from designed behavior.
TwinEthos recommendation (not law)
Scope every AI response and session cache to the requesting user or tenant
Key every cache of model responses, prompts, embeddings, or per-user AI state by the requesting user or tenant, and check ownership on every read. Detect caches keyed only by prompt or content hash that serve user-specific data, and session or history stores read without an owner check.
guardrail.opint-scoped-response-cache · TwinEthos derivation — guardrail.opint-scoped-response-cache · jurisdictions: *
maturity: reviewed · confidence: high
Why: A cache is the cheapest latency win in an AI stack and an easy way to hand one user's data to another. OpenAI has disclosed a cache fault that showed users other users' chat-history titles and payment details, and researchers have reported prompt caches shared across all users that could leak information between tenants.