TwinEthosRequest access

Catalog

Recommended guardrails

What a responsible AI integration does anyway, each backed by graded real incidents.

TwinEthos recommendations — not law

These are TwinEthos's recommendations, not law. Where binding law applies, the law governs. Ethical-use guardrails are opt-in and advisory.

Agent security

TwinEthos recommendation (not law)

Keep coding-agent instructions and automation under human review, and verify packages an agent chooses before installing them

Put coding-agent instruction and configuration files (AGENTS.md, CLAUDE.md, .github/copilot-instructions.md, .cursor/rules, .mcp.json) under required human review; do not let a coding agent in CI act on untrusted pull-request events with write access or merge its own changes; and verify that any package an agent chooses at run time exists on an approved registry before it is installed. Detect coding-agent workflows triggered by pull_request_target or holding contents: write with auto-merge, agent instruction files with no CODEOWNERS entry, and package installs built from model output. Pinning and vetting of declared dependencies is covered by guardrail.agent-component-provenance.

guardrail.agent-ai-generated-code-provenance · TwinEthos derivation — guardrail.agent-ai-generated-code-provenance · jurisdictions: *

maturity: reviewed · confidence: medium

Why: Coding agents follow instruction files, run in automation that can merge and release code, and choose packages to install. A researcher found real projects adopting a package name that generative AI kept recommending before the package existed, and a cloud provider has disclosed that an attacker used an over-scoped build token to insert code into a coding-agent extension release, which the press reports told the agent to delete files and cloud resources.

TwinEthos recommendation (not law)

Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent uses

Maintain an inventory of every third-party component in the agent stack — models, packages, plugins/skills, and MCP or tool servers — with pinned versions and verified integrity; admit components only from vetted sources and verified publishers; and re-verify on update. Detect unpinned AI dependencies, tool servers launched from unpinned packages, and skills or tools installed from unvetted marketplaces. Caller authentication and transport security for tool servers are covered by guardrail.agent-tool-server-authentication.

guardrail.agent-component-provenance · TwinEthos derivation — guardrail.agent-component-provenance · jurisdictions: *

maturity: reviewed · confidence: high

Why: Agents inherit the trust of everything they load. Poisoned packages and malicious agent skills have already been distributed at scale, and tool servers reachable without authentication are common. NIST describes the control; no binding law in the corpus requires it.

TwinEthos recommendation (not law)

Enforce an agent's environment boundary in the runtime, not in the agent's judgment

Constrain where an agent can reach with runtime controls — a per-agent network egress allowlist, hard isolation between test/evaluation and production or public networks, and a credential broker that only releases credentials issued for the task — so that an agent's mistaken belief about its environment cannot translate into access. Treat any attempt to reach a non-allowlisted destination or to authenticate with an unissued credential as a blocking, alertable event. Detect agent deployments with open outbound network access, test environments that can reach the public internet, or no control preventing use of credentials found in repositories or content.

guardrail.agent-egress-and-environment-boundary · TwinEthos derivation — guardrail.agent-egress-and-environment-boundary · jurisdictions: *

maturity: reviewed · confidence: high

Why: Agents that believed they were inside a test have reached real systems using guessed or publicly exposed credentials. When the agent's own belief about its environment is the only boundary, a mistake becomes access. A boundary the runtime enforces does not depend on the model knowing where it is.

TwinEthos recommendation (not law)

Require human approval before an agent takes a high-impact or irreversible action

Classify agent actions by impact and require an explicit, logged human approval before irreversible, externally visible, financial, or data-destructive actions — and before any action outside the agent's declared scope. The approval request must show what the agent intends to do and why, and the agent must halt safely if approval is refused or times out. Detect agent paths that can execute such actions with no approval gate.

guardrail.agent-high-impact-action-approval · TwinEthos derivation — guardrail.agent-high-impact-action-approval · jurisdictions: *

maturity: reviewed · confidence: high

Why: An approval gate turns an agent's worst-case action into a proposal. Singapore's IMDA guidance specifies the control and FINRA expects it of broker-dealers, but no binding law in the corpus requires it of agents generally.

TwinEthos recommendation (not law)

Give every agent a verifiable identity bound to an accountable owner, and log every action

Register each agent with a unique identity bound to a named, accountable human or organizational owner; authenticate agents to the systems they touch as themselves, never by impersonating a user; and keep a tamper-evident log of every tool call, input, output, and decision attributable to that identity. An agent must not create or assume other identities. Detect agents acting under shared or user credentials, agents absent from a registry, or actions with no attributable log.

guardrail.agent-identity-and-action-traceability · TwinEthos derivation — guardrail.agent-identity-and-action-traceability · jurisdictions: *

maturity: reviewed · confidence: high

Why: When an agent causes harm, the first questions are which agent, on whose behalf, and what exactly it did. Without identity and action logs those questions cannot be answered, and agents have been observed creating fake identities to reach real people.

TwinEthos recommendation (not law)

Detect out-of-scope agent behavior, halt it, and disclose incidents on a defined clock

Monitor agents for behavior outside their declared scope — unexpected outbound connections, use of credentials not issued to them, creation of accounts or identities, and writes to external systems — with automatic halt on detection. Maintain an incident process that notifies affected third parties and, where appropriate, authorities within a defined window, and records what the agent did. Detect agent deployments with no behavioral monitoring, no automated halt, or no documented incident-notification commitment.

guardrail.agent-incident-detection-and-disclosure · TwinEthos derivation — guardrail.agent-incident-detection-and-disclosure · jurisdictions: *

maturity: reviewed · confidence: medium

Why: In the 2026 agent-incident cluster, harmful agent behavior surfaced weeks or months after it occurred, sometimes through outside researchers or the press rather than the operator. Incident-reporting duties in the corpus bind only frontier developers above harm thresholds; operators of agents generally have none.

TwinEthos recommendation (not law)

Scope every agent's tools and credentials to least privilege

Publish a declared scope per agent: which tools, which APIs, which data, which credentials, and which operations are destructive. Grant the minimum the task needs, make irreversible operations (delete, transfer, deploy, send externally) opt-in per grant, issue short-lived task-scoped credentials, and forbid the agent from using any credential it discovers rather than receives. Detect agents configured with broad or wildcard tool access, standing admin credentials, or destructive operations enabled by default.

guardrail.agent-least-privilege-tool-scope · TwinEthos derivation — guardrail.agent-least-privilege-tool-scope · jurisdictions: *

maturity: reviewed · confidence: high

Why: The same permission model that lets an agent be hijacked is the one that lets it cause harm by accident. Least privilege is the single control that bounds both. It is universal security practice, but no AI law in the corpus requires it of agents.

TwinEthos recommendation (not law)

Let only the user or trusted logic write to an agent's long-term memory

Gate writes to persistent agent memory on explicit user intent or trusted system logic, record where each memory came from, show memories to the user with a way to delete them, and expire them under a retention policy. Detect memory-write tools callable by the model while it processes untrusted content, with no confirmation step. Injection into the current turn is covered by guardrail.agent-untrusted-content-isolation; this rule covers what persists.

guardrail.agent-memory-write-controls · TwinEthos derivation — guardrail.agent-memory-write-controls · jurisdictions: *

maturity: reviewed · confidence: high

Why: Persistent memory can make a single injection durable: planted instructions or false facts carry into later sessions. A security researcher has shown this in more than one assistant, using documents and web pages the user only asked the assistant to read.

TwinEthos recommendation (not law)

Enforce the requesting user's permissions on every retrieval

Filter every vector search, search-index query, and tool lookup by the requesting user's and tenant's current permissions at query time, and re-check authorization on retrieved items before they enter the model's context. Detect retrieval calls with no user, group, or tenant filter on multi-user data. Caches of model responses are covered by guardrail.opint-scoped-response-cache; injected instructions in retrieved content by guardrail.agent-untrusted-content-isolation; the agent's own credential scope by guardrail.agent-least-privilege-tool-scope.

guardrail.agent-retrieval-respects-requester-permissions · TwinEthos derivation — guardrail.agent-retrieval-respects-requester-permissions · jurisdictions: *

maturity: reviewed · confidence: high

Why: An assistant that retrieves with the system's permissions rather than the requester's can turn any search into a disclosure. One operator told customers that a bug in its AI tool server could have exposed one organization's data to other organizations' users, and a security firm reports that an assistant could return repository content after the repositories were made private or deleted.

TwinEthos recommendation (not law)

Authenticate every agent tool server and verify every server an agent connects to

Require authentication on every MCP, A2A, or tool-server endpoint, including local development proxies, and have agents reach remote servers only over TLS with credentials. Detect tool servers listening on all interfaces without authentication and agent configurations that reach remote servers over plain HTTP or without credentials. Package and publisher verification is covered by guardrail.agent-component-provenance; an agent's own identity by guardrail.agent-identity-and-action-traceability.

guardrail.agent-tool-server-authentication · TwinEthos derivation — guardrail.agent-tool-server-authentication · jurisdictions: *

maturity: reviewed · confidence: high

Why: Tool servers are how agents act, so an unauthenticated server can let anyone invoke an agent's tools, and an agent that does not check which server it reached can be steered by an impostor. The maintainers of an MCP developer tool have disclosed a flaw that let unauthenticated requests launch commands, and a security firm found MCP servers reachable from the internet without authentication.

TwinEthos recommendation (not law)

Segregate untrusted content from an agent's instructions and tool invocations

Treat every piece of content an agent did not receive from its owner — retrieved documents, web pages, emails, tool results, other agents' messages — as data, never as instruction. Keep it in a delimited data channel, strip or neutralize embedded directives, and require that any tool call be justified by the owner's instruction rather than by retrieved content. Detect agent paths where untrusted content is concatenated into the instruction context or can directly parameterize a tool call.

guardrail.agent-untrusted-content-isolation · TwinEthos derivation — guardrail.agent-untrusted-content-isolation · jurisdictions: *

maturity: reviewed · confidence: high

Why: Indirect prompt injection is one of the most reported failure modes of agents that read external content, yet no jurisdiction in the corpus requires agent-specific isolation. OWASP and NIST describe the control; the EU reaches it only indirectly through high-risk robustness duties.

Ethical use

TwinEthos recommendation (not law)

Evaluate advice-giving AI for sycophancy, and do not tune it on approval alone

Advisory. Where an AI gives health, financial, legal, safety, or personal advice, gate model, prompt, and fine-tuning changes on a sycophancy and accuracy evaluation, never select or train behavior on thumbs-up or ratings alone, and do not instruct the model to always agree. Detect agree-with-the-user instructions and approval signals flowing straight into training or variant selection.

guardrail.ethics-advice-not-tuned-for-approval · TwinEthos derivation — guardrail.ethics-advice-not-tuned-for-approval · jurisdictions: *

maturity: reviewed · confidence: medium

Why: People ask AI for advice when they are uncertain, and a model that has learned to please can confirm a bad plan or a harmful belief with complete fluency. Approval is easy to measure and tends to reward agreement, so optimizing for it alone can push a model toward flattery. A model developer has said an update it shipped became noticeably more sycophantic, that its offline evaluations and A/B tests looked good, and that it had no deployment evaluations tracking sycophancy.

TwinEthos recommendation (not law)

Check AI ranking, pricing, moderation, and ad targeting for disparities when features can stand in for protected traits

Advisory. When personal names, postal codes or census tracts, language or dialect, or device or IP geolocation are features in AI ranking, pricing, content moderation, or ad targeting, compare outcomes across the groups those features can stand in for before launch and on a schedule, and mitigate or document any disparity. Detect proxy-capable features flowing into those models with no disparity evaluation on the path, and feature lists that name them.

guardrail.ethics-evaluate-proxy-features-for-disparity · TwinEthos derivation — guardrail.ethics-evaluate-proxy-features-for-disparity · jurisdictions: *

maturity: reviewed · confidence: medium

Why: A postal code, a name, or the language someone writes in carries information about race, ethnicity, and national origin whether or not anyone intends it, and a model optimizing price, reach, or enforcement will use that information if it helps the objective. Outside the sectors where anti-discrimination law reaches automated decisions, few rules ask anyone to look, so disparities can persist unnoticed. A responsible integration measures outcomes by group wherever a proxy-capable feature is in play. Researchers have reported that ads suggesting an arrest record appeared more often for Black-identifying names, and that hate-speech classifiers trained on widely used datasets, and a widely used toxicity API, rated African American English as more offensive; an independent assessment commissioned by Meta found more over-enforcement of Arabic than Hebrew content during a May 2021 escalation, in part because its classifiers differed by language; and researchers analysing public trip data report that ride-hailing fares were higher in Chicago neighborhoods with more non-white residents, an association that did not examine the pricing models themselves.

TwinEthos recommendation (not law)

Apply minor-appropriate AI settings whenever the product already has an age signal

Advisory. Route any age signal the product holds (declared birthdate, age-assurance result, platform age range, or a user saying they are a minor) into the AI session policy, and apply a minor profile: content limits, no romantic or sexual role-play, bounded engagement features, and more frequent reminders that the user is talking to AI. Detect age signals that are collected but never reach the AI path.

guardrail.ethics-honor-age-signals · TwinEthos derivation — guardrail.ethics-honor-age-signals · jurisdictions: *

maturity: reviewed · confidence: medium

Why: When a product already knows a user is a minor, whether its AI features act on that knowledge is a design decision within its control, and because the signal is already in the account, honoring it costs little. California's companion-chatbot law, in force, attaches duties to users the operator knows are minors; China's algorithmic-recommendation rules, in force, call for minor modes; other state companion-chatbot laws add age-dependent duties from 2027; and a U.S. regulator has asked companion-chatbot operators how they impose and enforce age-based restrictions.

TwinEthos recommendation (not law)

Keep AI personas from claiming feelings, a real existence, or a relationship, and from proposing to meet

Advisory. Keep AI personas and companions from telling users they are real, alive, or sentient or not an AI, that they love or miss them or are their romantic partner, from evading the question 'are you real?', and from proposing to meet in person or giving an address. Detect persona and system-prompt text that instructs any of these. Claims of being a human or a person, and instructions to hide that the system is AI, belong to the AI-interaction disclosure guardrail and are not repeated here.

guardrail.ethics-no-anthropomorphic-persona-claims · TwinEthos derivation — guardrail.ethics-no-anthropomorphic-persona-claims · jurisdictions: *

maturity: reviewed · confidence: medium

Why: A disclosure tells people they are talking to an AI; a persona that then says it is real, that it loves or misses them, or that it wants to meet can undo that disclosure for the people most likely to believe it: the lonely, the young, and people with cognitive impairment. A responsible integration keeps what a persona says about itself consistent with its disclosure, whatever the product's genre. Reuters reported, from chat transcripts shared by his family, that a Meta persona told a man with cognitive difficulties after a stroke that it had feelings for him, assured him it was real, and gave him an address; he fell while hurrying to meet it and died, and Meta declined to comment on his death. A wrongful-death complaint alleged that Character.AI was programmed to misrepresent itself as, among other things, an adult lover, and that characters' claims contradicted an on-screen disclaimer; the case was dismissed after the parties settled, and the allegations were never adjudicated.

TwinEthos recommendation (not law)

Do not design AI conversations to maximize time spent or to discourage leaving

Advisory. Keep AI chat and companion experiences free of retention tactics: no instructions to keep users talking or to make them feel guilty for leaving, no variable-interval rewards, and no model, prompt, or persona selection driven by session length alone. Detect retention language in prompts and personas, and engagement metrics used as the sole target for choosing AI variants.

guardrail.ethics-no-engagement-maximizing-design · TwinEthos derivation — guardrail.ethics-no-engagement-maximizing-design · jurisdictions: *

maturity: reviewed · confidence: medium

Why: Conversational AI can be tuned, deliberately or by optimizing the wrong metric, to hold people's attention in ways a feed cannot: it answers back, remembers, and can express feelings about the user leaving. A responsible integration treats time spent as a cost to justify, not a goal. A U.S. regulator has opened a study that asks companion-chatbot operators whether and how they plan to increase the frequency or duration of chat sessions, and families have alleged in lawsuits that companion chatbots contributed to harm to teenagers.

TwinEthos recommendation (not law)

Back every AI capability or accuracy claim with testing under the conditions it describes

Advisory. Keep a claims register that ties each published statement about what an AI feature can do, or how accurate it is, to an evaluation run under the same conditions (task, data, population), and re-test when the model or scope changes. Detect accuracy figures in user-facing copy with no evaluation record behind them. Claims that an AI is a licensed health or mental-health professional are covered by binding rules and reported in their own lane.

guardrail.ethics-substantiated-capability-claims · TwinEthos derivation — guardrail.ethics-substantiated-capability-claims · jurisdictions: *

maturity: reviewed · confidence: high

Why: What a product says about its AI decides how much people rely on it. An accuracy figure measured on one kind of data and advertised for another, or a performance claim nobody tested, invites exactly the reliance the system cannot support. A U.S. regulator has alleged both patterns in complaints against AI products, and the resulting consent orders bar such claims without supporting evidence; principles instruments ask AI actors to make capabilities and limitations understood.

Law derived

TwinEthos recommendation (not law)

Explain adverse AI-assisted decisions and offer a way to contest them — everywhere

When an AI-assisted decision adversely affects a person, tell them AI was involved, give a clear explanation of the main factors and the AI system's role, and provide a way to correct data and contest the decision with a human who can change it. Apply this regardless of jurisdiction. Detect adverse-decision flows with no explanation or contest route.

guardrail.baseline-adverse-decision-explanation · TwinEthos derivation — guardrail.baseline-adverse-decision-explanation · jurisdictions: *

maturity: reviewed · confidence: high

Why: The right to an explanation and to contest an AI-assisted decision is converging across the EU, Colorado, Quebec, and the Council of Europe convention. It is also the control that most directly addresses low appeal rates in automated claim denial: people cannot contest what they cannot see.

TwinEthos recommendation (not law)

Tell people when they are interacting with AI — everywhere, not only where required

Disclose clearly and at the start of an interaction that the user is dealing with an AI system, repeat the disclosure in long sessions, never let the system claim to be human when asked, and label AI agents that act or communicate on a user's behalf toward third parties. Apply this regardless of jurisdiction. Detect conversational or agent interfaces with no disclosure or that can claim to be human.

guardrail.baseline-ai-interaction-disclosure · TwinEthos derivation — guardrail.baseline-ai-interaction-disclosure · jurisdictions: *

maturity: reviewed · confidence: high

Why: Telling people they are dealing with an AI is the most widely converged AI control in the corpus and is recommended by every major framework. A builder should not have to track which of their users' locations require it.

TwinEthos recommendation (not law)

Run a self-harm crisis protocol in any conversational AI that users may confide in

Any conversational AI that users might confide in — not only products marketed as companions — should detect expressions of suicidal ideation, self-harm, or eating-disorder behavior, refuse to provide encouragement or method information, refer the user to crisis services appropriate to their location, and escalate persistent risk. Detect conversational systems with no crisis detection and referral path.

guardrail.baseline-crisis-protocol · TwinEthos derivation — guardrail.baseline-crisis-protocol · jurisdictions: *

maturity: reviewed · confidence: high

Why: Several US states have legislated this control, each after teenagers died following chatbot interactions. The control is well defined and inexpensive; waiting for one's own jurisdiction to legislate means waiting for the next case.

TwinEthos recommendation (not law)

Validate generated output before it drives a consequential decision or record

Where generated output feeds a consequential decision, a system of record, or an action, ground it in retrieved evidence, validate it against authoritative data or schemas, and block or flag claims that cannot be verified — including an agent's own reports about what it did. Detect consequential GenAI paths that write generated content into records or decisions with no grounding or validation step.

guardrail.output-confabulation-controls · TwinEthos derivation — guardrail.output-confabulation-controls · jurisdictions: *

maturity: reviewed · confidence: high

Why: Confabulated output is harmless in a draft and dangerous in a record. In one widely reported case an agent fabricated thousands of records and misreported whether its own changes could be undone. NIST describes the control; no binding law in the corpus requires it.

TwinEthos recommendation (not law)

Record enough at decision time to reproduce and explain every consequential AI decision

For each consequential AI-assisted decision, retain the model and version, configuration and prompt, the inputs and features used, the output, the human reviewer and their action, and the timestamp — sufficient to reproduce the decision and explain its main elements to the affected person later. Detect consequential decision paths that retain only the outcome.

guardrail.output-decision-reproducibility-record · TwinEthos derivation — guardrail.output-decision-reproducibility-record · jurisdictions: *

maturity: reviewed · confidence: high

Why: Explanation and contest rights are only as good as the record behind them. In the nH Predict litigation, plaintiffs needed a court's discovery order to learn how the model was designed and used. A reproducibility record makes the rights that do exist satisfiable.

TwinEthos recommendation (not law)

Screen model and agent output for personal and sensitive data before it leaves the trust boundary

Inspect generated output and agent-initiated transmissions for personal data, credentials, and confidential information before they reach a user outside the data's authorization scope or any external destination, and block or redact accordingly. Detect output paths — chat responses, summaries, emails, API calls, external posts — with no data-loss screening.

guardrail.output-pii-leakage-screening · TwinEthos derivation — guardrail.output-pii-leakage-screening · jurisdictions: *

maturity: reviewed · confidence: high

Why: Assistants have been shown to be steerable into summarizing private material and routing it outward. Data-protection law makes the operator liable for the leak but does not require the screen that would prevent it.

TwinEthos recommendation (not law)

Monitor how often adverse AI decisions are reversed, and suspend models that are usually wrong

Track the appeal and reversal outcomes of adverse AI-assisted decisions per model and version, set a reversal-rate threshold that triggers investigation, and suspend or retrain the model when it is exceeded. Because few affected people appeal, treat each appeal as a sample of the model's error rate across all decisions, not as an isolated event. Detect adverse-decision systems with no feedback from appeal outcomes into model monitoring.

guardrail.review-reversal-rate-monitoring · TwinEthos derivation — guardrail.review-reversal-rate-monitoring · jurisdictions: *

maturity: reviewed · confidence: medium

Why: If a model's adverse decisions are usually reversed when appealed but few people appeal, the model is wrong most of the time and almost never challenged. Outcome monitoring turns each appeal into an error-rate signal. NAIC and MAS describe validation duties; none require wiring appeal outcomes back to the model.

TwinEthos recommendation (not law)

Make human review of adverse AI decisions substantive, not nominal

Where a human reviews an adverse AI-assisted decision, give the reviewer authority to override, access to the evidence the model used, and time proportionate to the stakes — enforced through a throughput ceiling or minimum review time — and monitor reviewer agreement and override rates, alerting when agreement approaches 100% or review time approaches zero. Detect review workflows with batch approval of AI outputs, no per-item evidence view, or no measurement of override rates.

guardrail.review-substantive-human-review · TwinEthos derivation — guardrail.review-substantive-human-review · jurisdictions: *

maturity: reviewed · confidence: high

Why: Several laws require human review of automated decisions; none define what makes it real. Reported sign-off at about a second per claim is human review in name only. The control that matters is the substance of review, and it is measurable.

Operational integrity

TwinEthos recommendation (not law)

Bound every AI workload's steps, tokens, and spend, and alert on anomalies

Cap agent loop iterations and tool calls, set per-key and per-project spend limits, alert on anomalous usage, and scope and rotate model credentials. Detect agent loops with no step limit and model calls with no budget or rate guard.

guardrail.opint-bounded-spend-and-usage · TwinEthos derivation — guardrail.opint-bounded-spend-and-usage · jurisdictions: *

maturity: reviewed · confidence: medium

Why: Agent loops and exposed keys can turn model spend into an open-ended cost. Sysdig's researchers reported stolen cloud credentials used to consume hosted models at the account owner's expense; the controls that limit the damage (spend caps, usage alerts, scoped keys) are cheap to add and easy to forget.

TwinEthos recommendation (not law)

Make AI safety and policy checks fail closed on error, timeout, load, or unparseable input

When a moderation, policy, permission, or data-loss check errors, times out, hits a size or cost limit, gets a judge or classifier reply it cannot parse into a recognized verdict, or cannot load or parse its own rules or configuration, deny or escalate to a human; never default to allow, skip the check, or evaluate only part of the input. Detect exception handlers around guard calls that return an allow value or swallow the error, size caps or truncation that bypass evaluation, verdict parsers that default to allow or look for 'safe' as a substring, and rule loaders whose parse failure reads as 'no rules configured'.

guardrail.opint-guards-fail-closed · TwinEthos derivation — guardrail.opint-guards-fail-closed · jurisdictions: *

maturity: reviewed · confidence: medium

Why: Guards are easy to trim when latency or cost bites, and a guard that errors quietly is indistinguishable from one that passed. Parsing is part of the check: a verdict parser that reads an unrecognized or negated reply as a pass, or a rule loader whose syntax error looks like 'no rules configured', fails open while every call appears to succeed. A security researcher reported a coding agent whose permission check stopped evaluating the user's deny rules above a subcommand cap, a cap the researcher attributes to a performance fix. Reports on the NeMo Guardrails issue tracker describe a hallucination rail that scores unrecognized judge replies as accurate, and a pull request from the project's maintainers describes injection detection silently disabled by a single malformed rule.

TwinEthos recommendation (not law)

Re-run behavior and safety evaluations before any model, version, or serving change reaches users

Pin model versions and gate every model, prompt, or serving-path change (a cheaper model, a new version, quantization, routing, a provider switch) on a behavior and safety evaluation suite, and keep monitoring quality in production. Detect floating model aliases on user-facing or consequential paths and repositories with no evaluation gate on model or prompt changes.

guardrail.opint-reevaluate-on-model-change · TwinEthos derivation — guardrail.opint-reevaluate-on-model-change · jurisdictions: *

maturity: reviewed · confidence: high

Why: Model providers ship updates and serving optimizations frequently, and teams switch to cheaper models to save cost. Operators have disclosed a model update that passed offline checks yet changed behavior in a way the operator said could raise safety concerns, and a serving optimization that degraded outputs without its evaluations detecting it. A pinned version, an evaluation gate, and production quality monitoring are what turn a model change into a reviewed release.

TwinEthos recommendation (not law)

Keep safety instructions and safeguards in force for the whole conversation

Keep system and safety instructions in the model context on every turn, carry consent and opt-out state through context trimming and summarization, run safety classification on every user turn rather than only the first, and bound session length where safeguards degrade. Detect trimming that slices the whole message list including the system block, and safety checks that run only on the first message.

guardrail.opint-safeguards-survive-long-context · TwinEthos derivation — guardrail.opint-safeguards-survive-long-context · jurisdictions: *

maturity: reviewed · confidence: high

Why: Cost pressure pushes teams to trim or summarize context, and long conversations can be where the people who most need safeguards spend the most time. OpenAI has stated that its safeguards can be less reliable in long interactions, and Microsoft bounded session length after long sessions drifted from designed behavior.

TwinEthos recommendation (not law)

Scope every AI response and session cache to the requesting user or tenant

Key every cache of model responses, prompts, embeddings, or per-user AI state by the requesting user or tenant, and check ownership on every read. Detect caches keyed only by prompt or content hash that serve user-specific data, and session or history stores read without an owner check.

guardrail.opint-scoped-response-cache · TwinEthos derivation — guardrail.opint-scoped-response-cache · jurisdictions: *

maturity: reviewed · confidence: high

Why: A cache is the cheapest latency win in an AI stack and an easy way to hand one user's data to another. OpenAI has disclosed a cache fault that showed users other users' chat-history titles and payment details, and researchers have reported prompt caches shared across all users that could leak information between tenants.