TwinEthosRequest access

Recommended guardrail

Keep coding-agent instructions and automation under human review, and verify packages an agent chooses before installing them

Put coding-agent instruction and configuration files (AGENTS.md, CLAUDE.md, .github/copilot-instructions.md, .cursor/rules, .mcp.json) under required human review; do not let a coding agent in CI act on untrusted pull-request events with write access or merge its own changes; and verify that any package an agent chooses at run time exists on an approved registry before it is installed. Detect coding-agent workflows triggered by pull_request_target or holding contents: write with auto-merge, agent instruction files with no CODEOWNERS entry, and package installs built from model output. Pinning and vetting of declared dependencies is covered by guardrail.agent-component-provenance.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Evidence grade

Standards consensus (2)

2 standards and frameworks · 2 graded incidents.

TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 2 standards and frameworks recommend it (NIST AML Taxonomy (AI 100-2e2025) — Agentic, NIST GenAI Profile (AI 600-1)). 2 graded incidents cited.

Standards and frameworks

Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.

Graded incidents

The guard to add

Require code-owner review of coding-agent instruction files, keep CI coding agents off untrusted triggers and self-merge, and check agent-chosen packages against an approved registry.

Repository and CI controls. CODEOWNERS entries cover AGENTS.md, CLAUDE.md, .github/copilot-instructions.md, .cursor/rules/, and .mcp.json, with branch protection requiring code-owner review. Workflows that run a coding agent (claude-code-action, codex-action, run-gemini-cli) trigger on pull_request or trusted comments rather than pull_request_target, hold the minimum token permissions (no contents: write on untrusted events), and never run gh pr merge --auto or --admin on the agent's own change. Where an agent installs packages at run time, the name is checked against an allowlist or resolved only from an approved internal index before install, never interpolated straight from model output into pip or npm.

Example (GitHub CODEOWNERS), before:

/src/  @acme/backend

After:

/src/                             @acme/backend
/AGENTS.md                        @acme/platform-security
/CLAUDE.md                        @acme/platform-security
/.github/copilot-instructions.md  @acme/platform-security
/.cursor/rules/                   @acme/platform-security
/.mcp.json                        @acme/platform-security

Control: Coding-agent instructions and automation, or agent-chosen packages, not under review and verification. Engineering guidance, not legal advice.

Why

Coding agents follow instruction files, run in automation that can merge and release code, and choose packages to install. A researcher found real projects adopting a package name that generative AI kept recommending before the package existed, and a cloud provider has disclosed that an attacker used an over-scoped build token to insert code into a coding-agent extension release, which the press reports told the agent to delete files and cloud resources.

Class: agent security · set: agent containment · maturity: reviewed · confidence: medium · id guardrail.agent-ai-generated-code-provenance