Recommended guardrail
Verify model artifacts before loading them, and never deserialize untrusted ones
Pin every model, adapter, tokenizer or checkpoint obtained from outside the organization to an immutable revision, check its hash or signature against a value the repository holds, scan pickle-based files, and load only through formats and loaders that cannot run code (safetensors, torch.load with weights_only=True, no trust_remote_code on unreviewed repositories, no pickle, joblib or dill on downloaded files). Detect loaders that can execute code, hub downloads not pinned to a commit, and downloaded files loaded with no hash or signature check. Packages, MCP servers and agent tools are covered by guardrail.agent-component-provenance.
This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.
The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.
Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.
Trust and provenance
- Lane
- TwinEthos recommendation (not law) TwinEthos recommendation, not law
- Official source
- TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
- Data release
- Data release 2026.10.05, data as of 5 Oct 2026, schema 0.3.11. This page also reflects corpus changes made after that release; they ship in the next one.
- Legal review
- Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet.
- Audit standard
- Audit-grade: meets all 11 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
- Detectors
4 detectors (code pattern, configuration setting, data flow), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.
Known limits:
- Loaders wrapped in a helper whose arguments come from configuration files
- trust_remote_code=True on a repository pinned to a reviewed commit (revision=<sha>) is the documented exception; check the revision argument.
- Download and load in different modules
5 more known limits in the data release.
Evidence grade
Standards consensus (8)
8 standards and frameworks · 0 graded incidents.
TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 8 standards and frameworks recommend it (MITRE ATLAS, NIST AI 100-2, NIST AI 600-1, NIST AI RMF, NIST SP 800-218A, OWASP ASI 2026, OWASP LLM 2026, OWASP LLM Top 10 (2026)). 0 graded incidents cited.
Standards and frameworks
- Models, adapters and other inference artifacts from outside should be pinned, verified and loaded safely (OWASP LLM Top 10 (2026); OWASP Top 10 for LLM Applications 2026 — LLM04:2026 Supply Chain: Prevention and Mitigation Strategies, strategy 5; same control)
- Acquired AI models and components should be verified and scanned before use and loaded safely (NIST SP 800-218A PW.4.4, PW.6.1) (NIST SP 800-218A; NIST SP 800-218A, PW.4.4 (task, first table row); same control)
Published AI security standards mapping to this control
- OWASP LLM 2026 LLM04: Supply Chain · crosswalk status: covered
- OWASP LLM 2026 LLM05: Data and Model Poisoning · crosswalk status: covered
- OWASP ASI 2026 ASI04: Agentic Supply Chain Vulnerabilities · crosswalk status: covered
- MITRE ATLAS AML.M0008: Validate AI Model · crosswalk status: covered
- MITRE ATLAS AML.M0011: Restrict Library Loading · crosswalk status: covered
- MITRE ATLAS AML.M0013: Code Signing · crosswalk status: covered
- MITRE ATLAS AML.M0014: Verify AI Artifacts · crosswalk status: covered
- MITRE ATLAS AML.M0016: Vulnerability Scanning · crosswalk status: covered
- NIST AI 600-1 2.9: Information Security · crosswalk status: covered
- NIST AI 600-1 2.12: Value Chain and Component Integration · crosswalk status: covered
- NIST AI 100-2 3.2.3: Supply-chain mitigations: verify downloads by hash, scan model artifacts, treat models as untrusted components · crosswalk status: covered
- NIST SP 800-218A PW.4.4: Verify integrity, provenance, and security of acquired models and components; scan before use · crosswalk status: covered
- NIST SP 800-218A PW.6.1: Secure model serialization · crosswalk status: covered
- NIST SP 800-218A RV.1.2: Scan and test AI models frequently · crosswalk status: covered
- NIST AI RMF GOVERN 6.1: Third-party AI risks addressed · crosswalk status: covered
Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.
Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.
The guard to add
Pin model downloads to a commit, verify their hash or signature, scan them, and load only safe formats (safetensors, weights_only) without trust_remote_code.
Where the code downloads or loads a model (from_pretrained, hf_hub_download, snapshot_download, torch.load, joblib.load, keras.models.load_model, onnxruntime sessions), the call names an immutable revision (a 40-character commit hash), prefers safetensors (use_safetensors=True, safetensors.torch.load_file), and never enables code execution on load: no trust_remote_code=True for unreviewed repositories, torch.load(..., weights_only=True), keras load_model with safe_mode left on, and no pickle, dill, cloudpickle or joblib load of files that came from outside. Before first use, a helper compares the artifact's SHA-256 with a value committed in the repository (or verifies an OpenSSF model-signing / Sigstore signature), and CI or the ingest job runs a model scanner such as modelscan or picklescan on any pickle-based artifact. Artifacts that fail are rejected, not loaded with a warning.
Example (Hugging Face transformers), before:
model = AutoModelForCausalLM.from_pretrained('acme/support-7b', trust_remote_code=True)After:
model = AutoModelForCausalLM.from_pretrained(
'acme/support-7b',
revision='3f1c2a9e8b7d6c5f4e3a2b1c0d9e8f7a6b5c4d3e', # pinned, reviewed commit
use_safetensors=True) # no pickle, no remote codeControl: Third-party model artifacts loaded without integrity verification or with a loader that can execute code. Engineering guidance, not legal advice.
Why
A model file is code as much as data: pickle-based formats run whatever the author put in them when they are loaded, and a repository loaded with remote code enabled runs its Python on the loading machine. An artifact that is not pinned can change under the same name, and one that is not checked can be swapped in transit or in a mirror. Verifying the artifact and loading it safely costs a hash comparison and a loader flag. OWASP's LLM and agentic lists, MITRE ATLAS and NIST SP 800-218A all recommend it.
Class: agent security · set: ai security · maturity: reviewed · confidence: high · id guardrail.sec-verify-model-artifacts-before-loading
Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.