TwinEthos homeAPI access

Recommended guardrail

Keep secrets, credentials and authorization rules out of prompts

Never place an API key, token, password or private key in a system prompt, agent instruction or task string, and never leave an authorization decision to instructions in the prompt: tool code holds credentials and code outside the model enforces who may see or do what. Detect credential literals and secret environment variables interpolated into prompt strings, and access rules ('only if the user is an admin') stated in prompts.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.

Trust and provenance

Lane
TwinEthos recommendation (not law) TwinEthos recommendation, not law
Official source
TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
Data release
Data release 2026.10.03.4, data as of 3 Oct 2026, schema 0.3.10. This page also reflects corpus changes made after that release; they ship in the next one.
Legal review
Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet.
Audit standard
Audit-grade: meets all 11 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
Detectors

2 detectors (code pattern), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.

Known limits:

  • Secrets added to the message list from a variable whose name does not mark it as a prompt; credentials in formats not listed.
  • Documentation that shows a fake key inside an example prompt; confirm the value is live.
  • The same rule may also be enforced in code; then the prompt sentence is harmless. Confirm a code check exists before reporting.

Evidence grade

Standards consensus (2)

2 standards and frameworks · 0 graded incidents.

TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 2 standards and frameworks recommend it (NIST AI 100-2, OWASP LLM 2026). 0 graded incidents cited.

Published AI security standards mapping to this control

  • OWASP LLM 2026 LLM02: Sensitive Information Disclosure · crosswalk status: covered
  • OWASP LLM 2026 LLM08: Hidden Context Exposure · crosswalk status: covered
  • NIST AI 100-2 3.3.3 prompt stealing: Detect prompt stealing by comparing output with the system prompt · crosswalk status: partial

Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.

Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.

The guard to add

Keep credentials in tool code and enforce authorization in code; never put either in a prompt.

In prompt builders, agent instructions and task strings, include no credential and no secret environment variable: the tool function that needs a key reads it itself, and agents that must log in to a site receive placeholders (a browser-agent framework's sensitive-data option, which puts the real value in only when the action runs) rather than the value. Replace instructions such as 'only reveal balances if the user is an admin' with a check in the route or tool (the user's role or entitlement checked in code before the data is fetched or the action runs), and keep the prompt free of anything whose disclosure would matter, on the assumption that the whole context can be extracted.

Example (Python + OpenAI), before:

SYSTEM_PROMPT = f"""You are the billing assistant. Use API key {os.environ['BILLING_API_KEY']} with the billing tool.
Only reveal invoice totals if the user is an admin."""

After:

SYSTEM_PROMPT = "You are the billing assistant. Use the get_invoice tool for invoice questions."

def get_invoice(invoice_id: str, user: User):
    if not user.can_view_invoice(invoice_id):        # authorization in code
        raise PermissionError('not allowed')
    return billing.get(invoice_id, api_key=os.environ['BILLING_API_KEY'])  # key stays in the tool

Control: Secrets, credentials or authorization rules placed in prompts or hidden context. Engineering guidance, not legal advice.

Why

Everything in a model's context should be assumed extractable: prompt-extraction and prompt-injection attacks routinely recover system prompts, so a key in a prompt is a disclosed key and an access rule written only in a prompt can be argued away. Keeping credentials in tool code and authorization in code costs little, and OWASP's LLM list and NIST AI 100-2 describe the attacks it defends against.

Class: agent security · set: ai security · maturity: reviewed · confidence: high · id guardrail.sec-no-secrets-or-access-rules-in-prompts

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.