TwinEthosRequest access

Control

User content used for model training without purpose-limited consent

User content (conversations, uploads, photos, voice) enters model training, fine-tuning, or evaluation datasets only with consent specific to that purpose, honoring withdrawal, and models trained on content without consent can be identified and retrained.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: AI design that manipulates, misleads, or neglects the people who use it · control id cond.user-content-trained-without-purpose-consent

Reach

1items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

The guard to add

Prefer checking a training-specific, unwithdrawn consent record before user content enters any training, fine-tuning, or evaluation dataset, and record lineage per model.

Consider a consent filter inside the dataset export job, the one place where conversations, uploads, photos, or voice clips become train.jsonl, a Hugging Face dataset, or preference pairs. It joins each item to a consent record whose purpose is model training (not general terms acceptance or service improvement) and drops items from users who never opted in or have withdrawn. The export writes a lineage manifest listing the content ids in each dataset and the dataset ids behind each training job, so a withdrawal removes the person's content from future runs and identifies models already trained on it for retraining.

Where it goes: 1 application source code, 2 data models, 11 CI/CD pipeline.

What reviewers look for: before any write of train.jsonl, Trainer(train_dataset=...), SFTTrainer or DPOTrainer, push_to_hub, or fine_tuning.jobs.create, a filter on a training-specific consent (consent.purpose == 'model_training' or user.training_opt_in) that excludes withdrawn users; a lineage record linking content to datasets and datasets to model versions.

Example (Python export + OpenAI fine-tuning), before:

rows = db.execute(text('SELECT id, user_id, prompt, reply FROM messages')).all()
write_jsonl('train.jsonl', [to_chat_example(r) for r in rows])
f = client.files.create(file=open('train.jsonl', 'rb'), purpose='fine-tune')
client.fine_tuning.jobs.create(training_file=f.id, model=BASE_MODEL)

After:

rows = db.execute(text("""
    SELECT m.id, m.user_id, m.prompt, m.reply FROM messages m
    JOIN consents c ON c.user_id = m.user_id
    WHERE c.purpose = 'model_training' AND c.granted AND c.withdrawn_at IS NULL""")).all()
write_jsonl('train.jsonl', [to_chat_example(r) for r in rows])
f = client.files.create(file=open('train.jsonl', 'rb'), purpose='fine-tune')
job = client.fine_tuning.jobs.create(training_file=f.id, model=BASE_MODEL)
lineage.record(job_id=job.id, dataset_file=f.id, content_ids=[r.id for r in rows])

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

TwinEthos recommendation (not law) (1)

Related incidents

  • FTC order requires Everalbum to delete face-recognition models trained on users' photos (2017-09; alleged (not proven)). The FTC alleged that Everalbum's Ever photo app enabled face recognition by default for most users and that, from September 2017 to August 2019, the company combined facial images extracted from users' photos with public datasets to develop its face-recognition technology, in part without affirmative express consent. Everalbum settled without admitting or denying the allegations; the final order (May 2021) requires deletion of face embeddings from users who had not consented and of any models or algorithms developed in whole or in part with Ever users' biometric information. Source: U.S. Federal Trade Commission (press release, 2021-05-07) · evidence grade: primary · cited by Train on user content only with consent specific to that purpose