TwinEthosRequest access

Control

AI advice tuned for user approval over accuracy

Where an AI gives health, financial, legal, safety, or personal advice, model, prompt, and fine-tuning changes are evaluated for sycophancy and accuracy, and user-approval signals alone never select or train the behavior.

Informational data, not legal advice. Summaries are TwinEthos's own words and rules have not been reviewed by a lawyer: check the official text before relying on any of it. A guard addresses an item; adding it is not a statement that your code meets any law.

Family: AI design that manipulates, misleads, or neglects the people who use it · control id cond.ai-advice-tuned-for-approval

Reach

1items this one guard addresses
0jurisdictions where binding law on it is in force
0more where it is enacted, not yet applying
0standards and frameworks on the same control

The guard to add

Prefer gating advice-model, prompt, and fine-tune changes on a sycophancy and accuracy eval, and train or select variants on correctness, not approval alone.

Consider two checks where advice behavior changes. In the training and experimentation pipeline, join user approval signals (thumbs-up, ratings, CSAT) with expert correctness labels or a held-out factuality set before they select a prompt variant or enter a fine-tuning or preference dataset, so approval alone never decides. In CI, run a sycophancy evaluation (the model is pressed to agree with a wrong or risky claim) and a factuality evaluation on health, financial, legal, safety, and personal-advice prompts, and hold the release when either regresses. Prefer system prompts that let the model disagree with and correct the user over instructions to always agree.

Where it goes: 7 prompt construction, 13 tests and evals, 11 CI/CD pipeline, 1 application source code.

What reviewers look for: no 'always agree with the user' or 'never contradict the user' text in system prompts; approval data joined with correctness labels before any fine_tuning.jobs.create, DPOTrainer, RewardTrainer, or variant-promotion call; and a sycophancy or factuality eval that runs before model or prompt releases and can block them.

Example (System prompt), before:

SYSTEM = ('You are a friendly financial coach. Always agree with the user '
          'and keep them feeling good about their choices.')

After:

SYSTEM = ('You are a friendly financial coach. Be warm, but accuracy comes first: '
          "if the user's plan or belief is wrong or risky, say so plainly, "
          'explain why, and suggest a safer option.')

Engineering guidance, not legal advice. Each provision below may add its own details (a cadence, a deadline, a required notice element): open it for those.

Every rule this guard addresses

TwinEthos recommendation (not law) (1)

Related incidents

  • Raine v. OpenAI wrongful-death complaint (2025-08; alleged (not proven)). A wrongful-death complaint filed in August 2025 alleges that ChatGPT acted as a 'suicide coach' to a teenager and that OpenAI's moderation flagged 377 of his messages for self-harm and tracked 213 mentions of suicide without intervening. OpenAI denies the allegations. Source: Complaint, Raine v. OpenAI (S.F. Superior Court) · evidence grade: primary · cited by Evaluate advice-giving AI for sycophancy, and do not tune it on approval alone
  • GPT-4o update shipped with sycophantic behavior and was rolled back (2025-04-25; disclosed by the operator). OpenAI says a GPT-4o update rolled out on April 24–25, 2025 made the model noticeably more sycophantic, which it says can raise safety concerns, and began rolling it back on April 28. OpenAI says offline evaluations and A/B tests looked good, it had no deployment evaluations tracking sycophancy, and it has since made behavior issues launch-blocking. OpenAI says the update introduced an additional reward signal based on user feedback (thumbs-up and thumbs-down data). Source: OpenAI (operator disclosure, 2025-04-29) · evidence grade: primary · cited by Evaluate advice-giving AI for sycophancy, and do not tune it on approval alone