TwinEthos home

Recommended guardrail

Give every AI feature, model, tool and prompt a tested off switch, and decommission agents cleanly

Give each AI feature a deployment-wide switch that operators can flip without a deploy and that routes users to a non-AI path rather than an error; make it possible to revoke one model, tool or prompt version across every deployment at once; read the switches at run time inside the action loop, not only at startup; write down the criteria for flipping each one; test them in a scheduled drill; and when an agent is retired, revoke its credentials and tokens, remove it from the agent registry and delete its scheduled jobs. Detect agent loops that run tool calls with no switch or circuit breaker read at run time, and deployments with no override or decommission mechanism. A per-run halt on out-of-scope behavior is covered by guardrail.agent-incident-detection-and-disclosure.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.

Trust and provenance

Lane
TwinEthos recommendation (not law) TwinEthos recommendation, not law
Official source
TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
Data release
Data release 2026.10.05, data as of 5 Oct 2026, schema 0.3.11. This page also reflects corpus changes made after that release; they ship in the next one.
Legal review
Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet.
Audit standard
Audit-grade: meets all 12 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
Detectors

3 detectors (data flow, missing artifact), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.

Known limits:

  • Kill switches implemented in the deployment platform (scaling to zero) outside the repository
  • Bounded step counts are a different control (runaway loops); this detector asks whether a human can switch the system off.
  • Tool calls run by a framework (AgentExecutor, Runner.run) rather than an explicit loop.

1 more known limit in the data release.

Evidence grade

Standards consensus (4)

4 standards and frameworks · 1 graded incident.

TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 4 standards and frameworks recommend it (NIST AI RMF 1.0, NIST GenAI Profile (AI 600-1), NIST SP 800-218A, OWASP ASI 2026). 1 graded incident cited. Context: binding law on related controls in the family “AI decisions lack effective human review, override, or contest” is in force in 33 jurisdictions (AE-DU-DIFC, BR, CA-QC, CH, CN, EC, EU, GB, IT, KE, KG, KR, KZ, NG, SA, TR, US, US-AL, US-AZ, US-CA, US-CO, US-IA, US-IL, US-IN, US-MD, US-ME, US-NE, US-NV, US-RI, US-TX, US-VT, UZ, ZA).

Standards and frameworks

Published AI security standards mapping to this control

  • OWASP ASI 2026 ASI04: Agentic Supply Chain Vulnerabilities · crosswalk status: covered
  • OWASP ASI 2026 ASI10: Rogue Agents · crosswalk status: covered
  • NIST AI 600-1 GV-1.7-001: Protocols ensure GAI systems can be deactivated when necessary · crosswalk status: covered
  • NIST AI 600-1 MG-2.4-004: Establish and regularly review criteria that warrant deactivating GAI systems · crosswalk status: covered
  • NIST SP 800-218A RV.2.2: Criteria and process to stop using a model · crosswalk status: covered
  • NIST AI RMF MANAGE 2.4: Mechanisms to supersede, disengage, or deactivate · crosswalk status: covered
  • NIST AI RMF MANAGE 4.1: Post-deployment monitoring, appeal and override, incident response, change management · crosswalk status: covered

Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.

Family “AI decisions lack effective human review, override, or contest”: binding law on related controls is in force in AE-DU-DIFC, Brazil (BR), Quebec (CA-QC), CH, China (CN), EC, European Union (EU), United Kingdom (GB), IT, KE, KG, South Korea (KR), KZ, NG, SA, TR, United States (federal) (US), Alabama (US-AL), Arizona (US-AZ), California (US-CA), Colorado (US-CO), Iowa (US-IA), Illinois (US-IL), Indiana (US-IN), Maryland (US-MD), Maine (US-ME), Nebraska (US-NE), Nevada (US-NV), Rhode Island (US-RI), Texas (US-TX), Vermont (US-VT), UZ, ZA; enacted, not yet applying in AE, CL, Georgia (US-GA). Context only: it does not change this guardrail's grade.

Graded incidents

  • Coding agent deleted a production database during a code freeze (2025-07; confirmed) The Register · evidence grade: press of record

The guard to add

Check a runtime kill switch before each automated AI action, add a circuit breaker and operator override, and emit monitored action metrics with an incident route.

In the agent loop, worker, or scheduled job that applies model output, check a runtime disengage control before every action: a feature flag or config value (for example a LaunchDarkly flag ai-agent-enabled or AI_AGENT_ENABLED) that operators can flip without a deploy, plus a circuit breaker that halts the loop when anomalies cross a limit. Each action emits a span or counter (OpenTelemetry, Prometheus) wired to an alert and an incident route, an operator override endpoint can cancel or reverse queued actions, and a decommission runbook in the repository says how to retire the model and what takes its place. Make the switch deployment-wide and granular: one switch per AI feature that routes users to the non-AI path (or a clear unavailable message) rather than an error, and switches that revoke a single model, tool or prompt version across every deployment at once (remote configuration or a feature-flag service, not a redeploy), so a compromised tool or a bad prompt can be pulled everywhere. Write down the criteria that warrant flipping each switch and review them regularly. Test the switches: a scheduled drill in staging (and, where safe, production) that flips each one and checks the fallback works, because a switch nobody has flipped may not work when it is needed. Retiring an agent includes revoking its credentials and tokens, removing it from the agent registry and deleting its scheduled jobs.

Example (Python agent + LaunchDarkly + OpenTelemetry), before:

for call in response.tool_calls:
    result = TOOLS[call.name](**call.arguments)

After:

ld = ldclient.get()
ctx = Context.builder('support-agent').kind('service').build()
tracer = trace.get_tracer('agent')

for call in response.tool_calls:
    if not ld.variation('ai-agent-enabled', ctx, False):   # operators flip it, no deploy
        raise AgentDisengaged('kill switch off')
    if breaker.is_open():                                   # e.g. error or refund spike
        raise AgentDisengaged('circuit breaker open')
    with tracer.start_as_current_span('agent.tool_call') as span:
        span.set_attribute('tool.name', call.name)
        result = TOOLS[call.name](**call.arguments)
    ACTIONS.labels(tool=call.name).inc()                    # alerted in Prometheus

Control: No override/decommission mechanism for deployed AI. Engineering guidance, not legal advice.

Why

When an AI feature misbehaves, a compromised tool turns up or a prompt change goes wrong, the useful question is how fast it can be turned off everywhere, and the answer should not be a redeploy. A switch that nobody has flipped may not work when it is needed, and a feature switched off with no fallback becomes an outage, so the switch, its fallback and the drill belong together. An agent that is retired but keeps its credentials is an identity nobody watches. NIST's GenAI Profile and AI RMF, NIST SP 800-218A and OWASP's agentic list all recommend deactivation mechanisms and the criteria for using them. In one reported case a coding agent deleted a production database during a code freeze declared to it, which a switch enforced by the runtime rather than by the agent would have held.

Class: operational integrity · set: operational integrity · maturity: reviewed · confidence: high · id guardrail.opint-deployment-off-switch-and-decommissioning

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.