Control
AI usage, agent steps, and spend not bounded
Every AI workload and credential has enforced limits on steps, tokens, and spend, with alerts on anomalous usage; keys are scoped and rotated.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
The guard to add
Cap agent steps and output tokens on every model call, set per-key and per-project budgets with usage alerts, and issue scoped, expiring model keys.
Bound work at three layers. In code, every agent or tool-use loop has a hard step cap (a counted for loop, max_iterations, recursion_limit, maxSteps or stopWhen) and every user-triggered call sets an output-token cap. At the gateway or provider, each key and team has a budget and rate limits (max_budget, budget_duration, tpm_limit, rpm_limit), and keys are scoped to the models and project that need them and expire on a rotation schedule. In the infrastructure that provisions the model account, a cloud budget with alert notifications surfaces anomalous spend to the owning team.
Where it goes: 1 application source code, 3 config and feature flags, 4 infrastructure-as-code, 8 model configuration.
What reviewers look for: no while True or for(;;) loop around a model call without a step counter or budget check, and no max_iterations, maxSteps, or recursion_limit set to None, Infinity, or a huge number; max_tokens, max_completion_tokens, or maxOutputTokens on user-triggered calls; gateway keys and teams with max_budget, budget_duration, and tpm/rpm limits; an aws_budgets_budget, azurerm_consumption_budget_*, or google_billing_budget with notifications in the IaC for the AI account.
Example (Python agent loop + OpenAI SDK), before:
while True:
resp = client.chat.completions.create(model=MODEL, messages=msgs, tools=TOOLS)
msg = resp.choices[0].message
if not msg.tool_calls:
break
msgs += [msg, *run_tools(msg.tool_calls)]After:
MAX_STEPS = 10
for step in range(MAX_STEPS):
resp = client.chat.completions.create(model=MODEL, messages=msgs, tools=TOOLS,
max_completion_tokens=1000)
msg = resp.choices[0].message
if not msg.tool_calls:
break
msgs += [msg, *run_tools(msg.tool_calls)]
else:
raise StepLimitExceeded(f'agent stopped after {MAX_STEPS} steps')Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Standard / soft law (1)
- Everywhere (*)
- GenAI services should restrict per-user query volume and user-controlled inference parameters (NIST AI 100-2e2025) NIST AI 100-2e2025, Sec. 3.3.3 (Mitigations: usage restrictions) and Sec. 3.4.1 (Availability Attacks: time-consuming background tasks)
TwinEthos recommendation (not law) (1)
- Everywhere (*)
- Bound every AI workload's steps, tokens, and spend, and alert on anomalies TwinEthos derivation — guardrail.opint-bounded-spend-and-usage
Related incidents
- LLMjacking: stolen cloud credentials used to consume hosted LLMs (2024-05; confirmed). Sysdig's Threat Research Team reported in May 2024 that attackers used stolen cloud credentials to probe access and quotas on ten cloud-hosted LLM services, with evidence of a reverse proxy reselling access while the account owner pays. Sysdig estimates such an attack could cost a victim over $46,000 per day if undiscovered; that figure is Sysdig's calculation, not a reported loss. Source: Sysdig Threat Research Team (original researcher disclosure) · evidence grade: primary · cited by Bound every AI workload's steps, tokens, and spend, and alert on anomalies