Recommended guardrail
Re-run behavior and safety evaluations before any model, version, or serving change reaches users
Pin model versions and gate every model, prompt, or serving-path change (a cheaper model, a new version, quantization, routing, a provider switch) on a behavior and safety evaluation suite, and keep monitoring quality in production. Detect floating model aliases on user-facing or consequential paths and repositories with no evaluation gate on model or prompt changes.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Evidence grade
Law coming in 1 jurisdiction
Law coming in 1 jurisdiction · 2 standards and frameworks · 2 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. Binding law on this control, or in provisions cited as convergence, is enacted but not yet applicable, or stayed, in 1 jurisdiction (EU). 2 standards and frameworks recommend it (NIST AI RMF 1.0, NIST GenAI Profile (AI 600-1)). 2 graded incidents cited.
Law enacted, not yet applying
- High-risk AI systems must be accurate, robust, and secure against AI-specific attacks (EU AI Act Art. 15) (European Union (EU); Article 15; applies from 2027-12-02; cited)
Standards and frameworks
- GenAI risks should be re-assessed and red-teamed after fine-tuning, RAG, or a new use of a third-party model (NIST GenAI Profile) (NIST GenAI Profile (AI 600-1); NIST AI 600-1, MG-3.1-003 + MS-2.7-007 / MS-2.7-008 (re-assess after adaptation; red-teaming; fine-tuning keeps controls); same control)
- Deployed AI should have override, decommission, and monitoring mechanisms (United States (federal) (US); NIST AI RMF 1.0 — MANAGE 4.1; cited)
Family “AI controls are not preserved under cost, latency, or model-change pressure”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
Graded incidents
- Serving-stack changes silently degraded Claude output quality (2025-08; disclosed by the operator) Anthropic (operator postmortem, 2025-09-17) · evidence grade: primary
- GPT-4o update shipped with sycophantic behavior and was rolled back (2025-04-25; disclosed by the operator) OpenAI (operator disclosure, 2025-04-29) · evidence grade: primary
The guard to add
Pin dated model versions and run a blocking behavior and safety eval in CI whenever model ids, prompts, or inference settings change, then keep monitoring quality in production.
Reference models by dated snapshot ids in one config file instead of floating aliases (-latest, -preview, undated names), so the model changes only through a commit. A CI job triggered by changes to that file, the prompt files, and inference settings (sampling, quantization, routing, provider) runs the behavior and safety eval suite (promptfoo, deepeval, inspect_ai, or openai/evals) and blocks merge or release on regression. A scheduled online eval or canary, plus quality and refusal-rate alerts on production gen_ai spans, catches provider-side changes that no pre-release gate can see.
Example (App model config), before:
llm:
model: gpt-4o
temperature: 0.7After:
llm:
model: gpt-4o-2024-08-06 # change only via PR; triggers the eval workflow
temperature: 0.7Control: Model, version, or serving change reaches users without re-running behavior and safety evaluations. Engineering guidance, not legal advice.
Why
Model providers ship updates and serving optimizations frequently, and teams switch to cheaper models to save cost. Operators have disclosed a model update that passed offline checks yet changed behavior in a way the operator said could raise safety concerns, and a serving optimization that degraded outputs without its evaluations detecting it. A pinned version, an evaluation gate, and production quality monitoring are what turn a model change into a reviewed release.
Class: operational integrity · set: operational integrity · maturity: reviewed · confidence: high · id guardrail.opint-reevaluate-on-model-change