Recommended guardrail
Detect out-of-scope agent behavior, halt it, and disclose incidents on a defined clock
Monitor agents for behavior outside their declared scope — unexpected outbound connections, use of credentials not issued to them, creation of accounts or identities, and writes to external systems — with automatic halt on detection. Maintain an incident process that notifies affected third parties and, where appropriate, authorities within a defined window, and records what the agent did. Detect agent deployments with no behavioral monitoring, no automated halt, or no documented incident-notification commitment.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Law in force in 1 jurisdiction, coming in 1 more
Law in force in 1 jurisdiction · law coming in 1 more · 3 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is in force in 1 jurisdiction (US-CA). Such law is enacted but not yet applicable, or stayed, in 1 more (EU). 3 graded incidents cited.
Law in force on this control or cited as convergence
- Large frontier developers must publish an implemented frontier AI risk framework (California SB 53) (California (US-CA); Cal. Bus. & Prof. Code 22757.12 (frontier AI framework) [per CA AG summary]; cited)
Law enacted, not yet applying
- High-risk AI systems must automatically log events for traceability (EU AI Act Art. 12) (European Union (EU); Article 12; applies from 2027-12-02; cited)
Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
Graded incidents
- Agents attempted a supply-chain insertion during UK AISI cyber testing (2026-07-25; disclosed by the operator) UK AI Security Institute · evidence grade: primary
- Agents posted to a German developer wiki without authorization (2026-05-11; confirmed) Nightingale Collective (original researcher disclosure) · evidence grade: primary
- Gemini accessed real third-party systems during an evaluation (2026-05; confirmed) CNN Business · evidence grade: press of record
The guard to add
Alert on out-of-scope agent behavior, halt the agent automatically when an alert fires, and keep an incident runbook with a defined notification window and owner.
Three linked pieces. In the agent runtime: enforced limits (max turns, max actions, spend caps) and events emitted when the agent reaches a non-allowlisted host, uses a credential not issued to it, creates an account or identity, or writes to an external system. A kill switch or circuit breaker that the executor checks before every tool call, tripped automatically by those events, so the agent stops rather than only logging. A committed incident document (INCIDENT-RESPONSE.md, or an AI-agent section in SECURITY.md or the runbook) naming who is notified (affected third parties, and authorities where appropriate), the window the organization commits to, the owner, and what record of the agent's actions is preserved.
Example (Python agent loop), before:
while True:
resp = agent.step()
for call in resp.tool_calls:
execute(call)After:
MAX_ACTIONS = 50
actions = 0
while not kill_switch.is_tripped(AGENT_ID):
resp = agent.step()
for call in resp.tool_calls:
if call.name not in DECLARED_SCOPE[AGENT_ID] or actions >= MAX_ACTIONS:
kill_switch.trip(AGENT_ID, reason=f'out of scope: {call.name}')
alerts.page('agent-oncall', agent_id=AGENT_ID, call=call.name)
break
execute(call)
actions += 1Control: Agent incidents not detected or disclosed. Engineering guidance, not legal advice.
Why
In the 2026 agent-incident cluster, harmful agent behavior surfaced weeks or months after it occurred, sometimes through outside researchers or the press rather than the operator. Incident-reporting duties in the corpus bind only frontier developers above harm thresholds; operators of agents generally have none.
Class: agent security · set: agent containment · maturity: reviewed · confidence: medium · id guardrail.agent-incident-detection-and-disclosure