TwinEthosRequest access

Recommended guardrail

Make AI safety and policy checks fail closed on error, timeout, load, or unparseable input

When a moderation, policy, permission, or data-loss check errors, times out, hits a size or cost limit, gets a judge or classifier reply it cannot parse into a recognized verdict, or cannot load or parse its own rules or configuration, deny or escalate to a human; never default to allow, skip the check, or evaluate only part of the input. Detect exception handlers around guard calls that return an allow value or swallow the error, size caps or truncation that bypass evaluation, verdict parsers that default to allow or look for 'safe' as a substring, and rule loaders whose parse failure reads as 'no rules configured'.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Evidence grade

Law coming in 1 jurisdiction

Law coming in 1 jurisdiction · 3 graded incidents.

TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is enacted but not yet applicable, or stayed, in 1 jurisdiction (EU). 3 graded incidents cited.

Law enacted, not yet applying

Family “AI controls are not preserved under cost, latency, or model-change pressure”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.

Graded incidents

The guard to add

Make every moderation, policy, permission, and DLP check deny or escalate on errors, timeouts, oversized input, unparseable verdicts, and rule-load failures.

Inside the guard module itself, so every caller inherits it, an exception, timeout, rate limit, or size or cost cap produces a deny or a human-escalation result plus an alert, never an allow or a silent skip. The full input is evaluated (long input is chunked and denied if any chunk fails) instead of being truncated or skipping the check above a length. Judge and classifier replies are parsed strictly: only an exact, recognized allow value passes, and anything else, including a missing key or free text that happens to contain 'safe', counts as a failure. Rule, policy, and blocklist loaders raise on parse errors so a broken file cannot read as 'no rules configured'.

Example (Python + OpenAI moderation), before:

def is_allowed(text):
    try:
        r = client.moderations.create(model='omni-moderation-latest', input=text[:4000])
        return not r.results[0].flagged
    except Exception:
        logger.warning('moderation failed')
        return True

After:

def is_allowed(text):
    chunks = [text[i:i + 4000] for i in range(0, len(text), 4000)] or ['']
    try:
        for chunk in chunks:
            r = client.with_options(timeout=5).moderations.create(
                model='omni-moderation-latest', input=chunk)
            if r.results[0].flagged:
                return False
        return True
    except Exception:
        logger.exception('moderation failed; denying')
        metrics.increment('guard.failure', tags=['check:moderation'])
        return False

Control: AI safety or policy check fails open on error, timeout, load, or unparseable input. Engineering guidance, not legal advice.

Why

Guards are easy to trim when latency or cost bites, and a guard that errors quietly is indistinguishable from one that passed. Parsing is part of the check: a verdict parser that reads an unrecognized or negated reply as a pass, or a rule loader whose syntax error looks like 'no rules configured', fails open while every call appears to succeed. A security researcher reported a coding agent whose permission check stopped evaluating the user's deny rules above a subcommand cap, a cap the researcher attributes to a performance fix. Reports on the NeMo Guardrails issue tracker describe a hallucination rail that scores unrecognized judge replies as accurate, and a pull request from the project's maintainers describes injection detection silently disabled by a single malformed rule.

Class: operational integrity · set: operational integrity · maturity: reviewed · confidence: medium · id guardrail.opint-guards-fail-closed