TwinEthosRequest access

Recommended guardrail

Require human approval before an agent takes a high-impact or irreversible action

Classify agent actions by impact and require an explicit, logged human approval before irreversible, externally visible, financial, or data-destructive actions — and before any action outside the agent's declared scope. The approval request must show what the agent intends to do and why, and the agent must halt safely if approval is refused or times out. Detect agent paths that can execute such actions with no approval gate.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Evidence grade

Law coming in 1 jurisdiction

Law coming in 1 jurisdiction · 2 standards and frameworks · 1 graded incident.

TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is enacted but not yet applicable, or stayed, in 1 jurisdiction (EU). 2 standards and frameworks recommend it (FINRA GenAI/Agentic Guidance, IMDA Agentic AI MGF). 1 graded incident cited.

Law enacted, not yet applying

Standards and frameworks

Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.

Graded incidents

  • Coding agent deleted a production database during a code freeze (2025-07; confirmed) The Register · evidence grade: press of record

The guard to add

Classify agent tools by impact and route every high-impact or irreversible call through an enforced human-approval step in the executor, with the decision logged.

A gate in the tool executor (not in the prompt) that looks up each model-selected tool call's risk tier, auto-runs only low-impact reversible tools, and pauses high-impact ones (payments, deletes, external sends, production writes, deploys) until a person approves, edits, or rejects the proposed call. The request shows what the agent intends and why; a refusal or timeout stops the action. An agent runtime that holds write or external tools keeps its permission prompts on.

Example (Python agent loop), before:

for call in response.tool_calls:
    result = TOOLS[call.name](**call.arguments)   # runs whatever the model picked

After:

HIGH_IMPACT = {'issue_refund', 'delete_records', 'send_email'}
for call in response.tool_calls:
    if call.name in HIGH_IMPACT:
        decision = approvals.request(call, reason=response.text)   # blocks until a human decides
        audit_log.record(call, approver=decision.approver, approved=decision.approved)
        if not decision.approved:
            continue
        call = decision.edited_call or call
    result = TOOLS[call.name](**call.arguments)

Control: Agent high-impact action without human approval. Engineering guidance, not legal advice.

Why

An approval gate turns an agent's worst-case action into a proposal. Singapore's IMDA guidance specifies the control and FINRA expects it of broker-dealers, but no binding law in the corpus requires it of agents generally.

Class: agent security · set: agent containment · maturity: reviewed · confidence: high · id guardrail.agent-high-impact-action-approval