subconscious.ai — fidelity report
Backend drift canary
2026-07-14

We re-ran five of our own replications. One of them got worse. Here is why.

On 2026-07-14 we re-ran five published conjoint studies on the current experiment backend (claude-sonnet-4 respondents) and graded them the hard way: level by level against each paper's reported human effects, and against our own best 2023–24 replication of the same study. Four of five produced scoreable panels. Small and mid-size designs held or improved. The hardest design regressed — and we traced every cause. This page is the drift canary we run before re-running the full replication corpus, published misses first.

Verdict

There is no uniform drift toward or away from human. Kreps2020 matched its best-ever original run (0.988 vs 0.989) and sits far above its 2023–24 median (0.747). Rapson2014 is stable and matches its July 2026 anchor exactly (0.941). Chaugule2015 slipped moderately (0.791 vs 0.983 original). But Hainmueller2014 — the hardest design on the page, 9 attributes and 50 parameters — fell to ρ = 0.45–0.68 across three runs, against 0.845 from its 2024 original and, more damningly, against 0.729 from the same pipeline and same model family in July 2026. That same-pipeline drop is the cleanest signal in the data. We traced it to experiment-design changes in our own backend — not to the model — and the fixes are committed below.

Headline numbers

Spearman ρ against the paper's human effects

StudyOriginal replication (best documented, 2023–24)Original distribution (all 2023–24 runs)July 2026 anchor (same pipeline)Current (2026-07-14, sonnet-4)Drift vs original
Hainmueller20140.845gpt-3.5-turbo-instruct, 2024-03-120.791n=11, 0.52–0.840.7290.512 / 0.676 / 0.447-0.17 regression
Rapson20141.000gpt-4o, 2024-10-181.000n=25, 0.00–1.000.9410.941 / 0.941-0.06 slightly down
Adida20190.872gemini-1.5-flash-002, 2024-11-010.858n=18, 0.59–0.96pending
Chaugule20150.983gpt4 (Azure), 2024-10-170.845n=12, 0.56–0.930.791 / 0.791-0.19 regression
Kreps20200.989gpt-4o, 2024-10-290.747n=9, 0.49–0.950.984 / 0.989stable

ρ computed identically for every arm: Spearman rank correlation between reported and synthetic effects over the matched level set, reference levels included (the pipeline's convention, and the convention behind the historical 0.55 / 0.73 numbers). Stricter non-reference-only values move every arm down together and do not change any verdict; they are in the methods section.

Calibration by study

Rank agreement, original arm vs current arm

Hainmueller2014

Original replication (2024-03-12, gpt-3.5-turbo-instruct) — ρ = 0.845
human effect ranksynthetic rank
Current backend (2026-07-14, sonnet-4) — ρ = 0.676
human effect ranksynthetic rank

Each dot is one attribute level from the published human study: x = its rank among the paper's reported effects, y = its rank among the synthetic effects. Perfect replication puts every dot on the diagonal. Original 2023–24 distribution for this study: median 0.791 across 11 runs (0.52–0.84).

Same-pipeline anchor (July 2026, sonnet-4): ρ = 0.729. That comparison holds the pipeline, estimator, and model family fixed.

Rapson2014

Original replication (2024-10-18, gpt-4o) — ρ = 1.000
human effect ranksynthetic rank
Current backend (2026-07-14, sonnet-4) — ρ = 0.941
human effect ranksynthetic rank

Each dot is one attribute level from the published human study: x = its rank among the paper's reported effects, y = its rank among the synthetic effects. Perfect replication puts every dot on the diagonal. Original 2023–24 distribution for this study: median 1.000 across 25 runs (0.00–1.00).

Same-pipeline anchor (July 2026, sonnet-4): ρ = 0.941. That comparison holds the pipeline, estimator, and model family fixed.

Adida2019

Original replication (2024-11-01, gemini-1.5-flash-002) — ρ = 0.872
human effect ranksynthetic rank
Current backend (2026-07-14, sonnet-4) — ρ = —

No data for this arm.

Each dot is one attribute level from the published human study: x = its rank among the paper's reported effects, y = its rank among the synthetic effects. Perfect replication puts every dot on the diagonal. Original 2023–24 distribution for this study: median 0.858 across 18 runs (0.59–0.96).

Chaugule2015

Original replication (2024-10-17, gpt4 (Azure)) — ρ = 0.983
human effect ranksynthetic rank
Current backend (2026-07-14, sonnet-4) — ρ = 0.791
human effect ranksynthetic rank

Each dot is one attribute level from the published human study: x = its rank among the paper's reported effects, y = its rank among the synthetic effects. Perfect replication puts every dot on the diagonal. Original 2023–24 distribution for this study: median 0.845 across 12 runs (0.56–0.93).

Kreps2020

Original replication (2024-10-29, gpt-4o) — ρ = 0.989
human effect ranksynthetic rank
Current backend (2026-07-14, sonnet-4) — ρ = 0.989
human effect ranksynthetic rank

Each dot is one attribute level from the published human study: x = its rank among the paper's reported effects, y = its rank among the synthetic effects. Perfect replication puts every dot on the diagonal. Original 2023–24 distribution for this study: median 0.747 across 9 runs (0.49–0.95).

Drift assessment

Reading the signal

  1. Capability on small designs is intact or better. Kreps2020 (7 attributes) at 0.984–0.988 ties the best run ever recorded for that study and sits far above the 2023–24 median of 0.747. On easy-to-mid designs the current backend performs at or above the original system's best, and well above its typical run.
  2. The regression is dimensional, not general. The one sharp drop is the study with 50 parameters. A 9-attribute, 3-alternative conjoint stresses exactly the things that change with a backend swap: prompt formatting of long profiles, order-effects, and attention over many attributes. Chaugule (13 params) slipped mildly; Kreps (18 params) didn't. The July anchor (0.729) is a single run, but the current number was replicated three times (0.512 / 0.676 / 0.447) — the depression is real, with high run-to-run variance on this design; the duplicate-run pairs on the other studies produced identical rank orders, so run-to-run rank noise is small on designs that work.
  3. The original arm is a biased-up comparator. "Original replication" here is the best-documented pre-2025 run per study; the full 2023–24 distributions (medians 0.75–1.00, minima as low as 0.0) show the old system was far noisier than its best runs suggest. Against distribution medians, the 2026-07-14 results win on two studies, tie on one, and lose on one. The right canary metric for the 150-study re-run is a paired distribution comparison, not best-vs-single.
  4. Answer to the question asked: the current backend is more human-aligned on three of four scored studies, and less on the highest-dimensional one. The drop was replicated (three runs) before publishing; the root cause is in the section below.

Root cause

The pipeline changed the experiments; the model was never the problem.

  1. Not the estimator. Re-scoring the 2024 run's raw choice data with the identical estimator gives 0.823 — the statistics explain none of the gap.
  2. Not the model. A same-day sweep across four model families (claude-sonnet-4, gpt-4o-mini, gpt-4o, llama-4) on the current backend all landed 0.38–0.78 with the same failure signature — while a far weaker 2023-era model scored 0.82–0.85 on the 2024 pipeline. When every model fails identically, the shared factor is the cause.
  3. The task framing was rewritten. The current backend regenerates the respondent instruction from the study topic instead of using the stored definition verbatim: respondents were asked to choose an "immigration policy option" instead of the paper's "which of the two immigrants would you prefer to see admitted." The stored definition was correct; the pipeline stopped honoring it.
  4. An opt-out was injected. The original study is a forced choice; the current backend adds a "no preference" escape. Synthetic respondents took it on 24–44% of the socially sensitive tasks (0% on value-neutral studies in the same batch), halving the effective sample.
  5. The instruction load grew ~22×. The 2024 task prompt was ~400 characters; the current one is ~10,000 — with injected mood, satisficing rules, deliberate response randomness, and a 13-statement attitude battery none of the replicated papers contain. Where the framing and opt-out bite, the socially loaded attributes collapse: the occupational-status gradient humans show (doctors strongly preferred) flattened to zero, exactly the social-desirability failure mode the critique literature predicts.

Status (updated 2026-07-16): seven backend defects and one client defect are filed with failing-test plans and exit criteria; a replication mode that runs stored definitions verbatim is in shaping, and a standing weekly fidelity canary — with a requested-vs-executed definition hash that catches any silent rewrite automatically — is specced. The intervention experiment (opt-out disabled; corrected framing; both) has not yet produced scoreable runs; the postscript below explains why. The full-corpus re-run is gated on those fixes.

Postscript — 2026-07-16

The canary caught a full backend outage the same afternoon.

Hours after the runs above completed, every new experiment on the same backend began failing at one mid-pipeline stage — the stage that generates the injected attitude battery documented above. That stage turned out to be hardcoded to a single external LLM provider, ignoring the model each experiment requested. When that provider's credential began rejecting requests, uncapped retry loops turned each run into a ~200,000-request storm (~87 requests/second for 38 minutes) until a timeout marked the run "finished" with zero output artifacts. The root cause was traced from the run logs and filed the same day. Two consequences matter for this page:

  1. A second silent substitution, now disclosed as a confound. On every current-arm run above, the injected attitude battery was generated by the hardcoded provider's model — not the model the experiment requested. The conjoint choices themselves used the requested model (its API traffic was healthy throughout), and the battery is injected noise in either case, so the conclusions stand — but it is exactly the class of silent substitution this page documents, found this time in the infrastructure rather than the prompt.
  2. The intervention arms are a miss so far. The first batch was killed by the outage. The resubmission that evening was accepted by the API and then silently dropped by the run queue — the same false-terminal status defect already on the filed list, observed from the other side. A minimal contract probe (two attributes × two levels, forced choice, small sample) now runs before any batch to confirm the pipeline executes end to end; the intervention results will be appended once a probe-gated batch completes and scores.

Methods and provenance

How every number was produced

Commitments

What happens before the full corpus re-run

  1. The backend gains a replication mode that executes stored study definitions verbatim: exact task wording, forced choice where the paper forced a choice, no injected battery or response randomness.
  2. The Hainmueller regression was replicated 3× before this page was published (0.512 / 0.676 / 0.447) — it is a real defect, not run noise, and it stays on this page until the fix lands and the number recovers.
  3. The full ~150-study re-run executes as paired distributions (every study × ≥2 replicates, this page's estimator) and compares medians — best-run-vs-best-run comparisons are banned.
  4. Results land here, misses included, with the same status labels as the rest of this site.