TwinEthos homeAPI access

Recommended guardrail

Authenticate self-hosted model, embedding and vector-store endpoints

Keep self-hosted inference, embedding and vector-store servers (Ollama, vLLM, llama.cpp server, Qdrant, Chroma, Weaviate) on loopback or a private network, or require an API key or authenticating proxy, and never enable anonymous access on a published port. Detect compose services that publish these ports on all interfaces with no authentication setting, start commands that bind them to 0.0.0.0 without an API key, Kubernetes LoadBalancer or NodePort Services in front of them, and Weaviate anonymous access. Agent tool servers are covered by guardrail.agent-tool-server-authentication.

TwinEthos recommendation — not law

This is TwinEthos's opinion of what a responsible AI integration does anyway. It is never a legal or standards requirement; where binding law applies, the law governs.

The recommended-guardrail rule files are open under CC BY 4.0; attribution and scope are in the terms.

Informational data, not legal advice. Summaries and rules have not been reviewed by a lawyer: always verify official law text for decisions. A suggested guard is intended to address each rule; adding it is not a statement of compliance to that law.

Trust and provenance

Lane
TwinEthos recommendation (not law) TwinEthos recommendation, not law
Official source
TwinEthos's own derivation record (from the corpus gap analysis and the incident registry), not an official source. The law, standards and incidents it cites are listed on this page with their own links.
Data release
Data release 2026.10.03.4, data as of 3 Oct 2026, schema 0.3.10. This page also reflects corpus changes made after that release; they ship in the next one.
Legal review
Not reviewed by a lawyer. Written by TwinEthos as its own recommendation: opinion, never law. No TwinEthos rule has been legally reviewed yet.
Audit standard
Audit-grade: meets all 11 checks of the TwinEthos audit standard that apply to it. The audit standard is TwinEthos's own quality bar for provenance, dates, applicability, detectors, fixtures, remediation and licences; it is not a legal review.
Detectors

1 detector (configuration setting), experimental: written from the rule's text and not yet measured for precision on real code, so treat a hit as a lead to verify.

Known limits:

  • Helm values files (a chart's service.type and API-key settings) are not inspected.
  • A ports: list placed before image: in the same compose service is not linked to the image.
  • A host firewall or private-network-only host may close the port; the compose file cannot show it.

Evidence grade

Standards consensus (2)

2 standards and frameworks · 0 graded incidents.

TwinEthos recommendation, not law. Where binding law applies, the law governs. No binding law in the corpus requires this control yet. 2 standards and frameworks recommend it (MITRE ATLAS, OWASP LLM 2026). 0 graded incidents cited.

Published AI security standards mapping to this control

  • OWASP LLM 2026 LLM06: Unbounded Consumption · crosswalk status: covered
  • OWASP LLM 2026 LLM09: Vector and Embedding Weaknesses · crosswalk status: partial
  • MITRE ATLAS AML.M0019: Control Access to AI Models and Data in Production · crosswalk status: covered

Item ids and titles from the published standards; the mapping is TwinEthos's (standards crosswalk, docs/COVERAGE.md Part 4). Cited by id, never quoted.

Family “An AI agent's authority, reach, inputs, and components are not bounded and accountable”: binding law on related controls is in force in no jurisdiction. Context only: it does not change this guardrail's grade.

The guard to add

Bind self-hosted model and vector endpoints to loopback or a private network, or put an API key in front of them.

In docker-compose, Kubernetes manifests and start scripts, publish inference and vector-store ports only on 127.0.0.1 ("127.0.0.1:11434:11434") or keep them on an internal network the application alone reaches. Where a port must be reachable, require authentication: vLLM serve --api-key (or VLLM_API_KEY), llama-server --api-key, QDRANT__SERVICE__API_KEY, a Chroma server auth provider (CHROMA_SERVER_AUTHN_PROVIDER), Weaviate API-key or OIDC authentication with AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED false, and for Ollama, which has no built-in authentication, an authenticating reverse proxy with the Ollama port itself unpublished.

Example (docker-compose), before:

services:
  ollama:
    image: ollama/ollama
    ports:
      - "11434:11434"
  qdrant:
    image: qdrant/qdrant
    ports:
      - "6333:6333"

After:

services:
  ollama:
    image: ollama/ollama
    ports:
      - "127.0.0.1:11434:11434"   # loopback only; the app reaches it on the compose network
  qdrant:
    image: qdrant/qdrant
    environment:
      QDRANT__SERVICE__API_KEY: ${QDRANT_API_KEY}
    ports:
      - "127.0.0.1:6333:6333"

Control: Self-hosted model, embedding or vector-store endpoint reachable without authentication. Engineering guidance, not legal advice.

Why

Self-hosted model and vector servers are often run from container quick-starts that publish the port on every interface with no credentials, and are then left that way in deployment. An open inference endpoint hands out the operator's compute and the model itself, and an open vector store exposes, and lets anyone alter, the documents behind every answer. OWASP's LLM list and MITRE ATLAS both recommend authenticating these endpoints.

Class: agent security · set: ai security · maturity: reviewed · confidence: high · id guardrail.sec-authenticated-model-and-vector-endpoints

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.