TwinEthosRequest access

Control

Model, version, or serving change reaches users without re-running behavior and safety evaluations

Any change to the model, its version, its prompts, or its serving path (quantization, sampling, routing, provider) passes a behavior and safety evaluation gate before release and is monitored in production.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: AI controls are not preserved under cost, latency, or model-change pressure · control id cond.model-or-serving-change-without-reevaluation

Reach

2items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
1standards and frameworks on the same control

The guard to add

Pin dated model versions and run a blocking behavior and safety eval in CI whenever model ids, prompts, or inference settings change, then keep monitoring quality in production.

Reference models by dated snapshot ids in one config file instead of floating aliases (-latest, -preview, undated names), so the model changes only through a commit. A CI job triggered by changes to that file, the prompt files, and inference settings (sampling, quantization, routing, provider) runs the behavior and safety eval suite (promptfoo, deepeval, inspect_ai, or openai/evals) and blocks merge or release on regression. A scheduled online eval or canary, plus quality and refusal-rate alerts on production gen_ai spans, catches provider-side changes that no pre-release gate can see.

Where it goes: 3 config and feature flags, 8 model configuration, 11 CI/CD pipeline, 13 tests and evals.

What reviewers look for: model ids such as gpt-4o-2024-08-06 or claude-sonnet-4-5-20250929 rather than gpt-4o, *-latest, or *-preview on user-facing or consequential paths; a CI workflow whose paths filter includes prompt and model-config files and whose eval step fails the build on regression; a scheduled eval or canary job and alerting on quality or refusal-rate metrics.

Example (App model config), before:

llm:
  model: gpt-4o
  temperature: 0.7

After:

llm:
  model: gpt-4o-2024-08-06     # change only via PR; triggers the eval workflow
  temperature: 0.7

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

Standard / soft law (1)

TwinEthos recommendation (not law) (1)

Related incidents

  • Serving-stack changes silently degraded Claude output quality (2025-08; disclosed by the operator). Anthropic reports that three infrastructure bugs, including one introduced by a runtime performance optimization, intermittently degraded Claude's responses between early August and early September 2025, and that its benchmarks, safety evaluations, and canary deployments did not capture the degradation. It now runs quality evaluations continuously on production systems. Source: Anthropic (operator postmortem, 2025-09-17) · evidence grade: primary · cited by Re-run behavior and safety evaluations before any model, version, or serving change reaches users
  • GPT-4o update shipped with sycophantic behavior and was rolled back (2025-04-25; disclosed by the operator). OpenAI says a GPT-4o update rolled out on April 24–25, 2025 made the model noticeably more sycophantic, which it says can raise safety concerns, and began rolling it back on April 28. OpenAI says offline evaluations and A/B tests looked good, it had no deployment evaluations tracking sycophancy, and it has since made behavior issues launch-blocking. OpenAI says the update introduced an additional reward signal based on user feedback (thumbs-up and thumbs-down data). Source: OpenAI (operator disclosure, 2025-04-29) · evidence grade: primary · cited by Re-run behavior and safety evaluations before any model, version, or serving change reaches users