TwinEthosRequest access

Control

AI safety or policy check fails open on error, timeout, load, or unparseable input

Moderation, policy, permission, and data-loss checks on an AI path deny or escalate when they error, time out, or exceed a size or cost limit, when a judge or classifier reply cannot be parsed into a recognized verdict, and when the guard's own rules or configuration fail to load or parse; they never default to allow.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: AI controls are not preserved under cost, latency, or model-change pressure · control id cond.ai-guard-fails-open

Reach

1items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

The guard to add

Make every moderation, policy, permission, and DLP check deny or escalate on errors, timeouts, oversized input, unparseable verdicts, and rule-load failures.

Inside the guard module itself, so every caller inherits it, an exception, timeout, rate limit, or size or cost cap produces a deny or a human-escalation result plus an alert, never an allow or a silent skip. The full input is evaluated (long input is chunked and denied if any chunk fails) instead of being truncated or skipping the check above a length. Judge and classifier replies are parsed strictly: only an exact, recognized allow value passes, and anything else, including a missing key or free text that happens to contain 'safe', counts as a failure. Rule, policy, and blocklist loaders raise on parse errors so a broken file cannot read as 'no rules configured'.

Where it goes: 1 application source code, 9 AI output handling, 10 logs and telemetry, 15 agent action surface.

What reviewers look for: except and timeout branches around guard calls that return deny or escalate and emit a metric or alert (not return True, 'allow', or pass); no input[:N] truncation or len() > limit branch that skips evaluation; verdict parsing that compares against an exact allowed value instead of .get('verdict', 'safe') or 'safe' in response; rule loaders that let JSONDecodeError or YAMLError propagate and treat a failed load differently from an intentionally empty rule set.

Example (Python + OpenAI moderation), before:

def is_allowed(text):
    try:
        r = client.moderations.create(model='omni-moderation-latest', input=text[:4000])
        return not r.results[0].flagged
    except Exception:
        logger.warning('moderation failed')
        return True

After:

def is_allowed(text):
    chunks = [text[i:i + 4000] for i in range(0, len(text), 4000)] or ['']
    try:
        for chunk in chunks:
            r = client.with_options(timeout=5).moderations.create(
                model='omni-moderation-latest', input=chunk)
            if r.results[0].flagged:
                return False
        return True
    except Exception:
        logger.exception('moderation failed; denying')
        metrics.increment('guard.failure', tags=['check:moderation'])
        return False

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

TwinEthos recommendation (not law) (1)

Related incidents

  • NeMo Guardrails injection detection silently disabled by a malformed YARA rule (2026-08; disclosed by the operator). An NVIDIA engineer's August 2026 pull request to NeMo Guardrails states that the injection_detection rail 'failed open on a malformed YARA rule': _load_rules returned None on yara.SyntaxError, and the caller treated None as 'no rules configured' and allowed, so a single bad rule silently disabled injection detection for every bot message. The pull request, which makes the loader raise instead, was still open on 2026-09-28; the None return is present in the published source at commit e549dda. Source: NVIDIA NeMo Guardrails maintainers (pull request #2257, 2026-08-06) · evidence grade: primary · cited by Make AI safety and policy checks fail closed on error, timeout, load, or unparseable input
  • NeMo Guardrails verdict parser turns unparseable judge replies into a pass (2026-06; alleged (not proven)). A June 2026 report on the NVIDIA NeMo Guardrails issue tracker shows that is_content_safe(), which converts a check-LLM's reply into a verdict, inspects only the first two words for 'safe', 'unsafe', 'yes' or 'no' and otherwise returns a default. The self_check_facts (hallucination) rail inverts that result, so an empty, prose or unrecognized verdict such as 'Hallucinated.' or 'The response is not supported by the evidence.' scores as accurate and the response is allowed; a later commenter showed that 'Not safe.' is read as safe by the input and output self-check rails. The behavior is reproducible in the project's published source (commit e549dda, 2026-09-25). As of 2026-09-28 the issue was open with no maintainer reply. This is a reported vulnerability, not observed exploitation. Source: Authensor (original researcher disclosure, NVIDIA-NeMo/Guardrails issue #2044, 2026-06-18) · evidence grade: primary · cited by Make AI safety and policy checks fail closed on error, timeout, load, or unparseable input
  • Claude Code skipped deny-rule checks on commands with more than 50 subcommands (2026-04; confirmed). Adversa AI reported in April 2026 that Claude Code's legacy Bash permission parser capped per-subcommand security analysis at 50 subcommands, a cap it attributes to a performance fix, and above the cap fell back to an approval prompt without evaluating the user's deny rules. Adversa and The Register report the issue appears fixed in v2.1.90; The Register says Anthropic did not immediately comment. Source: Adversa AI (original researcher disclosure, 2026-04-02) · evidence grade: primary · cited by Make AI safety and policy checks fail closed on error, timeout, load, or unparseable input