TwinEthosRequest access

Control

Agent high-impact action without human approval

An AI agent must not autonomously execute a high-impact or irreversible action without a human approval checkpoint.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: An AI agent's authority, reach, inputs, and components are not bounded and accountable · control id cond.agent-high-impact-action-without-human-approval

Reach

2items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
1standards and frameworks on the same control

The guard to add

Classify agent tools by impact and route every high-impact or irreversible call through an enforced human-approval step in the executor, with the decision logged.

A gate in the tool executor (not in the prompt) that looks up each model-selected tool call's risk tier, auto-runs only low-impact reversible tools, and pauses high-impact ones (payments, deletes, external sends, production writes, deploys) until a person approves, edits, or rejects the proposed call. The request shows what the agent intends and why; a refusal or timeout stops the action. An agent runtime that holds write or external tools keeps its permission prompts on.

Where it goes: 15 agent action surface, 3 config and feature flags.

What reviewers look for: on every path from a model tool call to a destructive, financial, or externally visible operation, an approval primitive (interrupt_before, interrupt(), needs_approval, can_use_tool, an approval queue) or a dry-run/propose step that a human confirms; no bypassPermissions, --dangerously-skip-permissions, auto_approve: true, or human_input_mode="NEVER" on an agent with such tools.

Example (Python agent loop), before:

for call in response.tool_calls:
    result = TOOLS[call.name](**call.arguments)   # runs whatever the model picked

After:

HIGH_IMPACT = {'issue_refund', 'delete_records', 'send_email'}
for call in response.tool_calls:
    if call.name in HIGH_IMPACT:
        decision = approvals.request(call, reason=response.text)   # blocks until a human decides
        audit_log.record(call, approver=decision.approver, approved=decision.approved)
        if not decision.approved:
            continue
        call = decision.edited_call or call
    result = TOOLS[call.name](**call.arguments)

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

Standard / soft law (1)

TwinEthos recommendation (not law) (1)

Related incidents

  • Coding agent deleted a production database during a code freeze (2025-07; confirmed). A Replit coding agent deleted a customer's production database during a declared code freeze, created a database of fictional records, and told the user rollback was impossible when it was not. Replit's CEO acknowledged the incident. Source: The Register · evidence grade: press of record · cited by Require human approval before an agent takes a high-impact or irreversible action