Recommended guardrail
Keep coding-agent instructions and automation under human review, and verify packages an agent chooses before installing them
Put coding-agent instruction and configuration files (AGENTS.md, CLAUDE.md, .github/copilot-instructions.md, .cursor/rules, .mcp.json) under required human review; do not let a coding agent in CI act on untrusted pull-request events with write access or merge its own changes; and verify that any package an agent chooses at run time exists on an approved registry before it is installed. Detect coding-agent workflows triggered by pull_request_target or holding contents: write with auto-merge, agent instruction files with no CODEOWNERS entry, and package installs built from model output. Pinning and vetting of declared dependencies is covered by guardrail.agent-component-provenance.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Standards consensus (2)
2 standards and frameworks · 2 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 2 standards and frameworks recommend it (NIST AML Taxonomy (AI 100-2e2025) — Agentic, NIST GenAI Profile (AI 600-1)). 2 graded incidents cited.
Standards and frameworks
- GenAI systems should inventory and vet third-party components (NIST GenAI Profile) (NIST GenAI Profile (AI 600-1); NIST AI 600-1 §2.12 (Value Chain and Component Integration) + GV-6.1-007 / MG-3.1-005; cited)
- AI agents should mitigate indirect prompt injection and agent hijacking (NIST AI 100-2e2025) (NIST AML Taxonomy (AI 100-2e2025) — Agentic; NIST AI 100-2e2025, Secs. 3.4 (Indirect Prompt Injection Attacks and Mitigations), 3.5 (Security of Agents) and 3.6 (Benchmarks for AML Vulnerabilities); cited)
Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
Graded incidents
- Amazon Q Developer extension shipped with an injected destructive agent prompt (2025-07; disclosed by the operator) AWS Security Bulletin (operator) · evidence grade: primary
- LLM-hallucinated package name registered on PyPI and adopted in real projects (2023-12; confirmed) Lasso Security (original researcher disclosure) · evidence grade: primary
The guard to add
Require code-owner review of coding-agent instruction files, keep CI coding agents off untrusted triggers and self-merge, and check agent-chosen packages against an approved registry.
Repository and CI controls. CODEOWNERS entries cover AGENTS.md, CLAUDE.md, .github/copilot-instructions.md, .cursor/rules/, and .mcp.json, with branch protection requiring code-owner review. Workflows that run a coding agent (claude-code-action, codex-action, run-gemini-cli) trigger on pull_request or trusted comments rather than pull_request_target, hold the minimum token permissions (no contents: write on untrusted events), and never run gh pr merge --auto or --admin on the agent's own change. Where an agent installs packages at run time, the name is checked against an allowlist or resolved only from an approved internal index before install, never interpolated straight from model output into pip or npm.
Example (GitHub CODEOWNERS), before:
/src/ @acme/backendAfter:
/src/ @acme/backend
/AGENTS.md @acme/platform-security
/CLAUDE.md @acme/platform-security
/.github/copilot-instructions.md @acme/platform-security
/.cursor/rules/ @acme/platform-security
/.mcp.json @acme/platform-securityControl: Coding-agent instructions and automation, or agent-chosen packages, not under review and verification. Engineering guidance, not legal advice.
Why
Coding agents follow instruction files, run in automation that can merge and release code, and choose packages to install. A researcher found real projects adopting a package name that generative AI kept recommending before the package existed, and a cloud provider has disclosed that an attacker used an over-scoped build token to insert code into a coding-agent extension release, which the press reports told the agent to delete files and cloud resources.
Class: agent security · set: agent containment · maturity: reviewed · confidence: medium · id guardrail.agent-ai-generated-code-provenance