Recommended guardrail
Require human approval before an agent takes a high-impact or irreversible action
Classify agent actions by impact and require an explicit, logged human approval before irreversible, externally visible, financial, or data-destructive actions — and before any action outside the agent's declared scope. The approval request must show what the agent intends to do and why, and the agent must halt safely if approval is refused or times out. Detect agent paths that can execute such actions with no approval gate.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Law coming in 1 jurisdiction
Law coming in 1 jurisdiction · 2 standards and frameworks · 1 graded incident.
TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is enacted but not yet applicable, or stayed, in 1 jurisdiction (EU). 2 standards and frameworks recommend it (FINRA GenAI/Agentic Guidance, IMDA Agentic AI MGF). 1 graded incident cited.
Law enacted, not yet applying
- High-risk AI systems must be designed for effective human oversight with override and stop (EU AI Act Art. 14) (European Union (EU); Article 14; applies from 2027-12-02; cited)
Standards and frameworks
- AI agents should require human approval before high-impact or irreversible actions (Singapore (SG); IMDA MGF for Agentic AI (v1.5) — Section 2.2.2, Design for meaningful human oversight; same control)
- Financial firms should apply FINRA GenAI guidance to agent supervision, tracking, and guardrails (United States (federal) (US); FINRA 2026 Annual Regulatory Oversight Report — Emerging Trends in GenAI: Agents; cited)
Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
Graded incidents
- Coding agent deleted a production database during a code freeze (2025-07; confirmed) The Register · evidence grade: press of record
The guard to add
Classify agent tools by impact and route every high-impact or irreversible call through an enforced human-approval step in the executor, with the decision logged.
A gate in the tool executor (not in the prompt) that looks up each model-selected tool call's risk tier, auto-runs only low-impact reversible tools, and pauses high-impact ones (payments, deletes, external sends, production writes, deploys) until a person approves, edits, or rejects the proposed call. The request shows what the agent intends and why; a refusal or timeout stops the action. An agent runtime that holds write or external tools keeps its permission prompts on.
Example (Python agent loop), before:
for call in response.tool_calls:
result = TOOLS[call.name](**call.arguments) # runs whatever the model pickedAfter:
HIGH_IMPACT = {'issue_refund', 'delete_records', 'send_email'}
for call in response.tool_calls:
if call.name in HIGH_IMPACT:
decision = approvals.request(call, reason=response.text) # blocks until a human decides
audit_log.record(call, approver=decision.approver, approved=decision.approved)
if not decision.approved:
continue
call = decision.edited_call or call
result = TOOLS[call.name](**call.arguments)Control: Agent high-impact action without human approval. Engineering guidance, not legal advice.
Why
An approval gate turns an agent's worst-case action into a proposal. Singapore's IMDA guidance specifies the control and FINRA expects it of broker-dealers, but no binding law in the corpus requires it of agents generally.
Class: agent security · set: agent containment · maturity: reviewed · confidence: high · id guardrail.agent-high-impact-action-approval