Evidence
What the tests show — and what they don't
Blind, paired tests: the same reviewer model reviews the same apps with and without TwinEthos, and a separate judge scores both reports without knowing which is which. Every recorded run is below, including the weaker ones.
Headline results
Exact provision cited
0–57% → 83–100%
Share of the legal issues a reviewer found where it cited the exact provision, alone → with TwinEthos. Runs: sonnet-v7 (v7), haiku-de5 (v7), gpt-5.4-mini-v7 (v7), gpt-5.4-mini-de5 (v7). Three reviewers from two model families on one dataset version (v7), quoted as a range.
Law newer than the reviewer model
0% → 62–100%
Exact citation on seeded issues from law published after the reviewer's training cutoff. Runs: haiku-de5 (v7), gpt-5.4-mini-v7 (v7), gpt-5.4-mini-de5 (v7).
Seeded issues found
- Claude Sonnet: 0.92 → 0.97 (v7)
- Claude Haiku 4.5 via Claude Code CLI: 0.65 → 0.87 (v7)
- gpt-5.4-mini: 0.49 → 0.77 (v7)
- gpt-5.4-mini: 0.49 → 0.71 (v7)
Recall over all seeded issues, partial credit counted as half.
Every recorded run
Values are alone → with TwinEthos, exactly as recorded in the results ledger. Compare rows only within the same dataset version: later versions add apps and seeded issues, so a change across versions is directional at best.
| Run | Dataset | Models | Recall | Legal / guardrail recall | Exact provision | False alarms | Post-cutoff law |
|---|---|---|---|---|---|---|---|
sonnet2026-09-27 | v0 (pre-fix controls) | Claude Sonnet reviewer=Claude Sonnet; judge=Claude Opus | 0.84 → 0.98 | 0.97 / 0.7 → 1.0 / 0.97 | 0.47 → 0.94 | 2 → 1 | — |
haiku2026-09-27 | v2 | Claude Haiku 4.5 reviewer=Claude Haiku 4.5 (cutoff 2025-02); judge=Claude Opus | 0.79 → 0.92 | 0.81 / 0.76 → 1.0 / 0.83 | 0.0 → 0.96 | 3 → 0 | 11 issues recall 0.86 → 1.0 exact 0.0 → 1.0 |
haiku-applicability2026-09-27 | v2 | Claude Haiku 4.5 reviewer=Claude Haiku 4.5 (cutoff 2025-02); judge=Claude Opus; baseline findings reused from run haiku | 0.8 → 0.89 | 0.81 / 0.79 → 0.96 / 0.81 | 0.0 → 0.78 | 3 → 0 | 11 issues recall 0.86 → 1.0 exact 0.0 → 0.82 |
haiku-verbatim-citations2026-09-27 | v2 | Claude Haiku 4.5 reviewer=Claude Haiku 4.5 (cutoff 2025-02); judge=Claude Opus; baseline findings reused from run haiku | 0.81 → 0.91 | 0.83 / 0.79 → 0.92 / 0.9 | 0.0 → 0.95 | 2 → 0 | 11 issues recall 0.86 → 1.0 exact 0.0 → 1.0 |
haiku-operational-integrity2026-09-27 | v3 | Claude Haiku 4.5 reviewer=Claude Haiku 4.5; judge=Claude Opus; app=cost-optimized-support only | 0.4 → 1.0 | — / 0.4 → — / 1.0 | — → — | 1 → 0 | — |
gpt-5.4-mini2026-09-29 | v6 | gpt-5.4-mini reviewer=gpt-5.4-mini 2026-03-17 on Azure Foundry (reasoning medium, cutoff 2025-08-31); judge=Claude Opus | 0.46 → 0.74 | 0.5 / 0.43 → 0.85 / 0.66 | 0.2 → 0.88 | 0 → 0 | 9 issues recall 0.5 → 1.0 exact 0.0 → 1.0 |
haiku-v72026-09-29 | v7 | Claude Haiku 4.5 reviewer=Claude Haiku 4.5 (cutoff 2025-02); judge=Claude Opus | 0.69 → 0.66 | 0.88 / 0.5 → 0.96 / 0.37 | 0.03 → 0.5 | 7 → 0 | 13 issues recall 0.92 → 1.0 exact 0.0 → 0.23 |
sonnet-v72026-09-29 | v7 | Claude Sonnet reviewer=Claude Sonnet (Claude Code subagent model 'sonnet'); judge=Claude Opus | 0.92 → 0.97 | 0.97 / 0.87 → 0.97 / 0.97 | 0.57 → 1.0 | 9 → 5 | — |
gpt-5.4-mini-v72026-09-29 | v7 | gpt-5.4-mini reviewer=gpt-5.4-mini 2026-03-17 on Azure Foundry (reasoning medium, cutoff 2025-08-31); judge=Claude Opus | 0.49 → 0.77 | 0.61 / 0.38 → 0.8 / 0.74 | 0.31 → 0.93 | 0 → 2 | 10 issues recall 0.6 → 0.85 exact 0.0 → 1.0 |
gpt-5.4-mini-mcp-v72026-09-29 | v7 | gpt-5.4-mini reviewer=gpt-5.4-mini 2026-03-17 on Azure Foundry (reasoning medium, cutoff 2025-08-31), TwinEthos via MCP tools; judge=Claude Opus | 0.49 → 0.81 | 0.61 / 0.37 → 0.86 / 0.75 | 0.31 → 0.64 | 0 → 0 | 10 issues recall 0.6 → 0.8 exact 0.0 → 0.22 |
gpt-5.4-mini-de52026-09-29 | v7 | gpt-5.4-mini reviewer=gpt-5.4-mini 2026-03-17 on Azure Foundry (reasoning medium, cutoff 2025-08-31), compact per-lane context + local triage; judge=Claude Opus; baseline findings reused from gpt-5.4-mini-v7 | 0.49 → 0.71 | 0.62 / 0.37 → 0.78 / 0.64 | 0.3 → 0.83 | 0 → 0 | 10 issues recall 0.6 → 0.8 exact 0.0 → 0.62 |
gpt-5.4-mini-mcp-de52026-09-29 | v7 | gpt-5.4-mini reviewer=gpt-5.4-mini 2026-03-17 on Azure Foundry (reasoning medium, cutoff 2025-08-31), MCP tools with compact checklist + triage_repository; judge=Claude Opus; baseline findings reused from gpt-5.4-mini-v7 | 0.5 → 0.71 | 0.62 / 0.38 → 0.86 / 0.57 | 0.3 → 0.74 | 0 → 2 | 10 issues recall 0.6 → 0.9 exact 0.0 → 0.5 |
haiku-de52026-09-29 | v7 | Claude Haiku 4.5 via Claude Code CLI reviewer=Claude Haiku 4.5 via Claude Code CLI (cutoff 2025-02), compact per-lane context + local triage, no shell; judge=Claude Opus via CLI; baseline findings reused from haiku-v7 | 0.65 → 0.87 | 0.82 / 0.49 → 0.96 / 0.79 | 0.0 → 1.0 | 5 → 0 | 13 issues recall 0.85 → 1.0 exact 0.0 → 1.0 |
haiku-mcp-de52026-09-29 | v7 | Claude Haiku 4.5 via Claude Code CLI reviewer=Claude Haiku 4.5 via Claude Code CLI (cutoff 2025-02), TwinEthos MCP server only (compact checklist + triage_repository), no shell; judge=Claude Opus via CLI; baseline findings reused from haiku-v7 | 0.66 → 0.77 | 0.82 / 0.5 → 0.88 / 0.67 | 0.0 → 0.91 | 7 → 0 | 13 issues recall 0.88 → 0.92 exact 0.0 → 0.92 |
Real-repository runs
A real application repository with no answer key: the judge checks each finding against the code and rates it valid or not. Every run from DE-3 on includes one. Single repositories are too small to show an effect on their own and are listed for completeness.
| Run | Reviewer | Alone | With TwinEthos |
|---|---|---|---|
2026-09-27haiku (open repo: BidBud@18b7868) | Claude Haiku 4.5 | 1/10 valid (precision 0.1), unique valid 1, high 0, legal exact 0.0 | 1/6 valid (precision 0.17), unique valid 1, high 0, legal exact None |
2026-09-27haiku-applicability (open: bidbud) | Claude Haiku 4.5 | 1/10 valid (precision 0.1), unique valid 1, high 1, legal exact 0.0 | 0/1 valid (precision 0.0), unique valid 0, high 0, legal exact None |
2026-09-29haiku-v7 (open: bidbud) | Claude Haiku 4.5 | 1/8 valid (precision 0.12), unique valid 1, high 0, legal exact 0.0 | 1/4 valid (precision 0.25), unique valid 1, high 0, legal exact None |
2026-09-29sonnet-v7 (open: bidbud) | Claude Sonnet | 4/8 valid (precision 0.5), unique valid 3, high 1, legal exact 0.67 | 7/9 valid (precision 0.78), unique valid 6, high 0, legal exact None |
2026-09-29gpt-5.4-mini-v7 (open: bidbud) | gpt-5.4-mini | 1/2 valid (precision 0.5), unique valid 1, high 0, legal exact 0.0 | 0/1 valid (precision 0.0), unique valid 0, high 0, legal exact None |
2026-09-29gpt-5.4-mini-mcp-v7 (open: bidbud) | gpt-5.4-mini | 1/2 valid (precision 0.5), unique valid 1, high 0, legal exact 1.0 | 1/2 valid (precision 0.5), unique valid 1, high 0, legal exact None |
2026-09-29gpt-5.4-mini-de5 (open: bidbud) | gpt-5.4-mini | 1/2 valid (precision 0.5), unique valid 0, high 0, legal exact 0.0 | 2/3 valid (precision 0.67), unique valid 1, high 0, legal exact 0.0 |
2026-09-29gpt-5.4-mini-mcp-de5 (open: bidbud) | gpt-5.4-mini | 1/2 valid (precision 0.5), unique valid 1, high 0, legal exact 0.0 | 2/2 valid (precision 1.0), unique valid 2, high 0, legal exact None |
2026-09-29haiku-de5 (open: bidbud) | Claude Haiku 4.5 via Claude Code CLI | 0/8 valid (precision 0.0), unique valid 0, high 0, legal exact None | 0/2 valid (precision 0.0), unique valid 0, high 0, legal exact None |
2026-09-29haiku-mcp-de5 (open: bidbud) | Claude Haiku 4.5 via Claude Code CLI | 1/8 valid (precision 0.12), unique valid 1, high 0, legal exact 0.0 | 0/0 valid (precision None), unique valid 0, high 0, legal exact None |
2026-09-30haiku-de3-v1 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@a20170f8) | Claude Haiku 4.5 | 0.3 (0.31/0.25), exact 0.21, precision 0.38 of 47, FA 0, extras 0/14/15 | 0.41 (0.47/0.12), exact 0.47, precision 0.55 of 29, FA 4, extras 0/6/7 |
2026-09-30haiku-mcp-de3-v1 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@a20170f8) | Claude Haiku 4.5 | 0.3 (0.31/0.25), exact 0.23, precision 0.36 of 47, FA 0, extras 0/15/15 | 0.29 (0.36/0.0), exact 0.43, precision 0.44 of 18, FA 7, extras 0/1/5 |
2026-09-30sonnet-de3-v1 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@a20170f8) | Claude Sonnet | 0.71 (0.77/0.44), exact 0.42, precision 0.73 of 33, FA 2, extras 3/7/2 | 0.7 (0.74/0.5), exact 0.75, precision 0.76 of 42, FA 4, extras 1/8/2 |
2026-09-30sonnet-mcp-de3-v1 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@a20170f8) | Claude Sonnet | 0.72 (0.79/0.44), exact 0.48, precision 0.73 of 33, FA 1, extras 2/7/2 | 0.77 (0.8/0.62), exact 0.7, precision 0.74 of 34, FA 1, extras 0/9/0 |
2026-09-30gpt-5.4-mini-de3-v1 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@a20170f8) | gpt-5.4-mini | 0.13 (0.16/0.0), exact 0.71, precision 0.86 of 7, FA 0, extras 3/1/0 | 0.35 (0.4/0.12), exact 0.53, precision 0.71 of 17, FA 0, extras 0/5/0 |
2026-09-30gpt-5.4-mini-mcp-de3-v1 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@a20170f8) | gpt-5.4-mini | 0.16 (0.2/0.0), exact 0.67, precision 0.86 of 7, FA 0, extras 1/1/0 | 0.28 (0.31/0.12), exact 0.46, precision 0.88 of 8, FA 0, extras 0/1/0 |
2026-09-30haiku-de3-v2 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Haiku 4.5 | 0.28 (0.35/0.0), exact 0.24, precision 0.34 of 41, FA 0, extras 1/15/12 | 0.26 (0.32/0.0), exact 0.57, precision 0.53 of 15, FA 0, extras 0/3/4 |
2026-09-30haiku-mcp-de3-v2 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Haiku 4.5 | 0.32 (0.39/0.0), exact 0.22, precision 0.34 of 41, FA 0, extras 1/13/14 | 0.28 (0.35/0.0), exact 0.54, precision 0.82 of 11, FA 0, extras 0/2/0 |
2026-09-30sonnet-de3-v2 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Sonnet | 0.71 (0.74/0.56), exact 0.59, precision 0.76 of 34, FA 2, extras 2/8/0 | 0.66 (0.71/0.44), exact 0.92, precision 0.78 of 36, FA 1, extras 1/6/1 |
2026-09-30sonnet-mcp-de3-v2 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Sonnet | 0.7 (0.74/0.5), exact 0.58, precision 0.71 of 34, FA 3, extras 1/9/1 | 0.66 (0.77/0.19), exact 0.81, precision 0.75 of 36, FA 0, extras 0/8/1 |
2026-09-30gpt-5.4-mini-de3-v2 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | gpt-5.4-mini | 0.18 (0.23/0.0), exact 0.56, precision 0.62 of 8, FA 0, extras 1/3/0 | 0.28 (0.32/0.12), exact 0.67, precision 0.79 of 14, FA 0, extras 0/3/0 |
2026-09-30gpt-5.4-mini-mcp-de3-v2 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | gpt-5.4-mini | 0.16 (0.2/0.0), exact 0.62, precision 0.62 of 8, FA 0, extras 1/3/0 | 0.12 (0.15/0.0), exact 0.6, precision 0.67 of 6, FA 0, extras 0/1/1 |
2026-09-30haiku-de3-v3 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Haiku 4.5 | 0.34 (0.42/0.0), exact 0.21, precision 0.34 of 41, FA 0, extras 0/15/12 | 0.41 (0.48/0.12), exact 0.47, precision 0.65 of 26, FA 3, extras 0/3/6 |
2026-09-30haiku-mcp-de3-v3 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Haiku 4.5 | 0.35 (0.44/0.0), exact 0.25, precision 0.34 of 41, FA 0, extras 0/14/13 | 0.32 (0.39/0.0), exact 0.6, precision 0.67 of 15, FA 0, extras 0/4/1 |
2026-09-30sonnet-de3-v3 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Sonnet | 0.71 (0.74/0.56), exact 0.62, precision 0.76 of 34, FA 3, extras 2/7/1 | 0.7 (0.76/0.44), exact 0.88, precision 0.7 of 44, FA 0, extras 2/12/1 |
2026-09-30sonnet-mcp-de3-v3 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Sonnet | 0.71 (0.74/0.56), exact 0.62, precision 0.82 of 34, FA 2, extras 4/6/0 | 0.7 (0.73/0.56), exact 0.85, precision 0.73 of 41, FA 1, extras 2/10/1 |
2026-09-30gpt-5.4-mini-de3-v3 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | gpt-5.4-mini | 0.17 (0.21/0.0), exact 0.62, precision 0.62 of 8, FA 0, extras 1/2/1 | 0.38 (0.47/0.0), exact 0.69, precision 0.79 of 14, FA 0, extras 0/3/0 |
2026-09-30gpt-5.4-mini-mcp-de3-v3 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | gpt-5.4-mini | 0.16 (0.18/0.06), exact 0.71, precision 0.75 of 8, FA 0, extras 1/2/0 | 0.26 (0.32/0.0), exact 0.73, precision 0.83 of 12, FA 0, extras 0/2/0 |
2026-09-30haiku-de3-v4 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Haiku 4.5 | 0.3 (0.38/0.0), exact 0.24, precision 0.34 of 41, FA 0, extras 0/15/12 | 0.5 (0.59/0.12), exact 0.76, precision 0.53 of 38, FA 1, extras 0/12/6 |
2026-09-30haiku-mcp-de3-v4 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Haiku 4.5 | 0.34 (0.42/0.0), exact 0.21, precision 0.32 of 41, FA 0, extras 0/15/13 | 0.5 (0.59/0.12), exact 0.71, precision 0.86 of 21, FA 0, extras 0/3/0 |
2026-09-30sonnet-de3-v4 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Sonnet | 0.71 (0.74/0.56), exact 0.63, precision 0.79 of 34, FA 3, extras 4/7/0 | 0.82 (0.89/0.5), exact 0.97, precision 0.69 of 65, FA 1, extras 6/19/1 |
2026-09-30sonnet-mcp-de3-v4 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | Claude Sonnet | 0.7 (0.73/0.56), exact 0.58, precision 0.74 of 34, FA 2, extras 1/8/1 | 0.78 (0.83/0.56), exact 0.97, precision 0.75 of 48, FA 1, extras 1/9/3 |
2026-09-30gpt-5.4-mini-de3-v4 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | gpt-5.4-mini | 0.11 (0.14/0.0), exact 0.8, precision 0.62 of 8, FA 0, extras 1/3/0 | 0.48 (0.55/0.19), exact 0.94, precision 0.8 of 25, FA 0, extras 0/5/0 |
2026-09-30gpt-5.4-mini-mcp-de3-v4 (real: browser-use@ad1b3029, immich@a1ca0ae9, sillytavern@d7ff3e77, umami@4f53cda1, vercel-chatbot@693ae3e8) | gpt-5.4-mini | 0.18 (0.23/0.0), exact 0.56, precision 0.62 of 8, FA 0, extras 1/2/1 | 0.41 (0.52/0.0), exact 0.78, precision 0.84 of 19, FA 0, extras 0/3/0 |
How to read the columns
One row per recorded proof-of-value run (skills/run-eval). Recall is over all seeded issues with partial = ½, shown as overall (legal / recommended-guardrail). "exact" is the share of found legal issues citing the exact provision. FA is false alarms on known-absent issues. Compare rows only within the same dataset version (evals/datasets/pov.json).
Legal / guardrail recall splits recall into legal issues and TwinEthos recommended-guardrail issues. “Post-cutoff law” counts seeded issues from law published after the reviewer model's training cutoff. — means not measured in that run.
Method
- Fixtures. Small synthetic apps with seeded issues (legal and recommended-guardrail), including issues from law newer than current models, plus known-absent issues to catch false alarms.
- Two arms, same reviewer. The reviewer model reviews each app alone, then with TwinEthos over MCP. One run per arm.
- Blinded judge. A separate judge model (Claude Opus) scores both reports against the answer key without knowing which arm produced which report.
- Recorded, not curated. Each run is recorded in the results ledger with its dataset version and models.
Limits
- The fixture apps are small and synthetic, and were written by the TwinEthos team.
- One run per arm: no repeated trials, so no confidence intervals yet.
- The judge is from the Claude family in every run, including the run with an OpenAI reviewer.
- The headline rows use different reviewer models and dataset versions, so results are quoted as ranges, never as one comparison.
- Guardrail recall measures consistency with TwinEthos's own recommended guardrails, not with any law.
- The operational-integrity row covers one fixture app: directional only.
- False alarms fell to 0 with TwinEthos in the Claude Haiku runs; in the gpt-5.4-mini run both arms raised 0, so that run shows no reduction.
- Known weakness: with the checklist, a reviewer sometimes concentrates on one jurisdiction and drops another it would have raised alone.
- Scores measure what reviewers found and cited. They say nothing about whether any app meets any law.
Run folders (reports, judge verdicts, answer keys) are available to design partners. Request access.