fidelity.subconscious.ai/study / Ditto evidence

AI-vs-human meta-study

A compact audit surface for the Ditto replication corpus: the study registry, estimator comparisons across generations, parent papers, and per-run W&B evidence.

New (2026-07-14): the backend drift canary — five replications re-run against our own 2023–24 results and the papers' human effects, one confirmed regression, root cause traced and published, misses first.

154
Runnable sources
json + published AMCE csv
150
Parent paper links
DOI / URL / eprint
128
W&B run-name rows
local replication ledger
126
2023-baselined studies
original replications with a labeled baseline
4
LLM replication rows
curated model-study matrix

Charts

Static charts generated from the same CSVs as the tables below.

Evidence Coverage Parent links 150 W&B run names 128 Drive links 0

Rows with usable parent paper links, W&B run-name evidence, or Drive links.

Model Replication Score Hainmueller2014 / databricks-claude-sonnet-4 0.686 Hainmueller2014 / databricks-llama4-maverick 0.640 Hainmueller2014 / azure-openai-gpt4o-mini 0.631 Hainmueller2014 / azure-openai-gpt4 0.630

Curated sample row from the model replication matrix.

Golden-10 paired local CLM Kreps2020 / 2024 0.893 Kreps2020 / current 0.897 Leng2021 / 2024 0.935 Leng2021 / current 0.919 Adida2019 / 2024 0.905 Adida2019 / current 0.857 Donnaloja2022 / 2024 0.486 Donnaloja2022 / current 0.413 Hainmueller2014 / 2024 0.706 Hainmueller2014 / current 0.465 Humburg2015 / 2024 0.353 Humburg2015 / current 0.726 Spilker2018 / 2024 0.398 Spilker2018 / current 0.427 Hainmueller2015_candidate / 2024 0.557 Hainmueller2015_candidate / current 0.022 Eliasson2017 / 2024 0.446 Eliasson2017 / current 0.767 Schweizer2012 / 2024 0.357 Schweizer2012 / current 0.679

Human-paper coefficient ranks compared with the same local CLM scorer: paper-era 2023-24 artifacts and current production.

Golden-10 production benchmark

Paired comparison against human-paper coefficients using the same local CLM estimator and alignment for both eras. 2024 local CLM mean: 0.604; Current local CLM mean: 0.617; paired mean delta: +0.014. Paper-era rows preserve their actual 2023-24 run dates.

RankStudyHuman N2024 local CLM ρCurrent local CLM ρDeltaPaper2024 runCurrent run
1Kreps202019710.8930.897+0.005Paper2024 W&BCurrent W&B
2Leng202118830.9350.919-0.017Paper2024 W&BCurrent W&B
3Adida201918000.9050.857-0.048Paper2024 W&BCurrent W&B
4Donnaloja202215970.4860.413-0.073Paper2024 W&BCurrent W&B
5Hainmueller201414070.7060.465-0.241Paper2024 W&BCurrent W&B
6Humburg20159030.3530.726+0.374Paper2024 W&BCurrent W&B
7Spilker20188000.3980.427+0.029Paper2024 W&BCurrent W&B
8Hainmueller2015_candidate3110.5570.022-0.535Paper2024 W&BCurrent W&B
9Eliasson20172850.4460.767+0.321Paper2024 W&BCurrent W&B
10Schweizer2012910.3570.679+0.321Paper2024 W&BCurrent W&B

LLM replication matrix

Verified model bake-off (2026-07-14, Hainmueller2014, k=5 replicates per model, 400 synthetic respondents each): ex-ref OLS Spearman vs the paper's published AMCEs. All four models land at 0.63–0.69; claude-sonnet-4's edge is run-to-run consistency (sd 0.031 vs 0.08–0.14), not mean. Rows are curated in data/model_replications.csv; reproduce with meta_study/scripts/score_runs_example.py.

StudyDomainMarketTransitionExperiment runModelHumanLLMScoreStatusW&B
Hainmueller2014Public Policymarket_us_public_policy_immigrationtransition_voter_prefers_immigrant_candidatebakeoff_2026_07_14databricks-claude-sonnet-4published AMCEs (41 effects)0.6860.686verifiedW&B
Hainmueller2014Public Policymarket_us_public_policy_immigrationtransition_voter_prefers_immigrant_candidatebakeoff_2026_07_14databricks-llama4-maverickpublished AMCEs (41 effects)0.6400.640verifiedW&B
Hainmueller2014Public Policymarket_us_public_policy_immigrationtransition_voter_prefers_immigrant_candidatebakeoff_2026_07_14azure-openai-gpt4o-minipublished AMCEs (41 effects)0.6310.631verifiedW&B
Hainmueller2014Public Policymarket_us_public_policy_immigrationtransition_voter_prefers_immigrant_candidatebakeoff_2026_07_14azure-openai-gpt4published AMCEs (41 effects)0.6300.630verifiedW&B

Estimator comparison (canary studies)

Same human baselines, three estimator generations. 2023 CLM = the ORIGINAL replications (inclusive convention, data/historic_2023_replications.csv). 2025 MNP = the small-n validation sweep (hb_core2, inclusive) — a different estimator generation; never conflate (see docs/verification/2026-07-21). Local OLS = client-side OLS-LPM on raw choices, ex-ref convention (Ditto#137). Server HB = the current hierarchical-Bayes AMCE tables, same convention — divergent pending rehoboam#2014/#2018. Ported MNL/CLM = rehoboam's own MLE and conditional-logit code copied into hb_validation (Ditto#159), run locally on the same labeled data, ex-ref convention. Rows curated in data/estimator_comparison.csv; the full old-era ledger (817 runs, 72 studies, pooled mean ρ 0.42) is committed as data/historic_replications.csv.

Study2023 CLM ρ (original)2025 MNP ρ (validation)Local OLS ρPorted MNL ρPorted CLM ρServer HB ρNotes
Hainmueller20140.5406 (n=1)0.413 (n=12)0.566 (n=2)0.7130.7540.202 (n=19)2026-07-14 bake-off, 4 models pooled; HB defect rehoboam#2014/#2018; 2026-07-17 ported estimators (n=3 sonnet-4 runs): MNL 0.675/0.735/0.729, CLM 0.774/0.849/0.638; golden-10 wave 07-19: 0.552/0.580 spread 0.028 PASSES but ~0.1 below the 07-14 band (0.65-0.73) - DRIFT SIGNAL, config diff pending
Kreps20200.6211 (n=5)0.872 (n=10)0.933 (n=2)0.917-0.008 (n=2)2026-07-17 canary drained: OLS 0.917/0.950 (mean 0.933) beats the 2023 original baseline (0.62) and the 2025 MNP sweep (0.87); HB 0.017/-0.033 on the same runs - strongest #2018 evidence; 2026-07-17 CLM 0.917 both runs; MNL n/a (41 unmapped level ids break the ported text-join); AUDIT 2026-07-17: scores cover 6/7 attributes - the Vaccine-efficacy attribute is unmapped (run logged continuous 50-90 vs level ids), design deviation; golden-10 07-19: k=1 failed_missing_artifacts (#2024), k=2 quarantined #2025 (0.90 on 6/7 attrs)
Hess20120.5004 (n=1)0.597 (n=10)2026-07-14 submissions lost in worker-pool stall (no runs materialized); covered by the 142-study batch run
Donnaloja20220.4381 (n=1)0.354 (n=7)2026-07-14 submissions lost in worker-pool stall (no runs materialized); covered by the 142-study batch run
Breitenstein20190.5213 (n=1)0.3995 (n=10)0.754 (n=2)0.7540.754golden-10 wave 2026-07-19: pair spread 0.000 (identical rank order both runs); PASSES stability; conformant
Rapson20140.7827 (n=1)0.1167 (n=10)0.90 (n=2)0.900.90golden-10 wave 07-19: clm paper, headline CLM 1.00/0.80 mean 0.90; spread 0.20 breaches 0.15 but = one rank swap on a ~5-rank design (granularity caveat); conformant; k=2 was a crash-replacement
Jung20180.1155 (n=1)0.7544 (n=10)0.655 (n=2)0.821golden-10 wave 07-19: UNSTABLE - ols spread 0.26 (0.524/0.786) breaches 0.15; choice-model estimators tighter (CLM 0.90/0.74, MNL 0.81/0.79); transcription triage queued (#121-adjacent)
Kikulwe20110.829 (n=21)0.3333 (n=10)-0.143golden-10 wave 07-19: k=1 clm paper headline -0.14 (all estimators agree, conformant run) - real miss, transcription triage first; k=2 finished w/o artifacts (rehoboam#2024)
Mele20210.1859 (n=8)golden-10 wave 07-19: quarantined #2025 (Age->continuous 39-51); 6-attr scores 0.80,1.00 recorded not quoted
Leng20210.5579 (n=2)golden-10 wave 07-19: quarantined #2025 (efficacy+adverse-rate continuous); noise 0.59/0.95
Spilker20180.1549 (n=5)golden-10 wave 07-19: quarantined #2025 (countries->continuous); 0.67/0.63 on corrupted design
Azimy20200.6119 (n=2)golden-10 wave 07-19: quarantined #2025; -0.38/-0.38

Study registry

Replication scope: the 142 studies with W&B objects (historic ledger runs or linked runs), out of 154 runnable conjoint experiments in the corpus. Each row links to its paper (DOI), its W&B runs (sign-in required), and its showcase page. 2023 ρ̄ is the per-study mean Spearman of the ORIGINAL 2023 replications (CLM pipeline, inclusive convention; data/historic_2023_replications.csv). The separate 2025-02 MNP validation sweep lives in the estimator table below; Best ρ is the current ex-ref OLS score where re-scored. The same data is machine-readable at /studies.json.