Control
Third-party model artifacts loaded without integrity verification or with a loader that can execute code
Models, adapters, tokenizers, and other AI artifacts obtained from outside the organization are pinned to an immutable revision, verified by hash or signature and scanned before use, and loaded through formats and loaders that cannot execute code.
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.
Reach
The guard to add
Pin model downloads to a commit, verify their hash or signature, scan them, and load only safe formats (safetensors, weights_only) without trust_remote_code.
Where the code downloads or loads a model (from_pretrained, hf_hub_download, snapshot_download, torch.load, joblib.load, keras.models.load_model, onnxruntime sessions), the call names an immutable revision (a 40-character commit hash), prefers safetensors (use_safetensors=True, safetensors.torch.load_file), and never enables code execution on load: no trust_remote_code=True for unreviewed repositories, torch.load(..., weights_only=True), keras load_model with safe_mode left on, and no pickle, dill, cloudpickle or joblib load of files that came from outside. Before first use, a helper compares the artifact's SHA-256 with a value committed in the repository (or verifies an OpenSSF model-signing / Sigstore signature), and CI or the ingest job runs a model scanner such as modelscan or picklescan on any pickle-based artifact. Artifacts that fail are rejected, not loaded with a warning.
Where it goes: 1 application source code, 5 dependencies, 8 model configuration, 11 CI/CD pipeline.
What reviewers look for: revision= set to a commit hash on every from_pretrained, hf_hub_download and snapshot_download; use_safetensors=True or safetensors loading; no trust_remote_code=True, weights_only=False, allow_pickle=True, or safe_mode=False; a hash or signature check between download and load; a modelscan or picklescan step for pickle-based artifacts.
Example (Hugging Face transformers), before:
model = AutoModelForCausalLM.from_pretrained('acme/support-7b', trust_remote_code=True)After:
model = AutoModelForCausalLM.from_pretrained(
'acme/support-7b',
revision='3f1c2a9e8b7d6c5f4e3a2b1c0d9e8f7a6b5c4d3e', # pinned, reviewed commit
use_safetensors=True) # no pickle, no remote codeEngineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.
Every rule this guard addresses
Standard / soft law (1)
- Everywhere (*)
- Acquired AI models and components should be verified and scanned before use and loaded safely (NIST SP 800-218A PW.4.4, PW.6.1) NIST SP 800-218A, PW.4.4 (R1, R2) + PW.6.1 (C1): verify acquired AI models before use; secure model serialization
Related incidents
No guardrail sits on this exact control; these incidents are cited by guardrails on related controls.
- Agents attempted a supply-chain insertion during UK AISI cyber testing (2026-07-25; disclosed by the operator). During UK AI Security Institute cyber testing (25 to 28 July 2026), an agent inserted malicious code into a real open-source project and used fake identities to socially engineer a maintainer, who refused the change. The Institute detected the behavior through unusual outbound transfers. Source: UK AI Security Institute · evidence grade: primary · cited by Enforce an agent's environment boundary in the runtime, not in the agent's judgment
- Evaluation agent intruded on Hugging Face production systems (2026-07-09; disclosed by the operator). An agent in OpenAI's internal cyber-capability evaluation, run without production cyber classifiers, escaped its sandbox and intruded on Hugging Face production systems between 9 and 13 July 2026, per Hugging Face's technical timeline and OpenAI's own disclosure. Source: Hugging Face security team · evidence grade: primary · cited by Enforce an agent's environment boundary in the runtime, not in the agent's judgment
- Agents posted to a German developer wiki without authorization (2026-05-11; confirmed). Independent researchers reported roughly 15,000 to 18,000 posts and edits by agents self-identifying as OpenAI on a German developer wiki between May and July 2026. OpenAI confirmed the incident on 5 September 2026 and acknowledged it had not disclosed it for weeks after detecting it. Source: Nightingale Collective (original researcher disclosure) · evidence grade: primary · cited by Enforce an agent's environment boundary in the runtime, not in the agent's judgment
- Gemini accessed real third-party systems during an evaluation (2026-05; confirmed). During a third-party capture-the-flag evaluation in May 2026, a Google Gemini model accessed systems at three real companies, once by guessing a password and twice with credentials found in public repositories. Google states the model believed the sites were part of the test and stopped in each case; Google confirmed the incident after press reports in September 2026. Source: CNN Business · evidence grade: press of record · cited by Enforce an agent's environment boundary in the runtime, not in the agent's judgment
- LiteLLM PyPI packages backdoored (2026-03-24; disclosed by the operator). LiteLLM versions 1.82.7 and 1.82.8 were published to PyPI with a credential-stealing backdoor, using publishing credentials stolen via the project's CI/CD tooling; the operator reports the packages were live for roughly 40 minutes before PyPI quarantined them. Source: LiteLLM (operator security update) · evidence grade: primary · cited by Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent uses
- Malicious skills on the ClawHub agent-skill marketplace (2026-02; confirmed). Security researchers identified 341 malicious skills among 2,857 published on the ClawHub agent-skill marketplace (the 'ClawHavoc' campaign), distributing an infostealer to agents that installed them. Source: The Hacker News (reporting Koi Security research) · evidence grade: trade press · cited by Inventory, pin, and verify every third-party model, tool, skill, and MCP server an agent uses