Experiments / expC_causal_verification
expC_causal_verification
Causal verification pipeline — correlational-to-causal survival rate
Result maps
Hypothesis
Hypothesis — expC run #1 (does the agreement probe map survive ablation?)
Causal verification pipeline — correlational→causal survival rate. Registered 2026-08-12 before the run. Input: the differential (real−twin) agreement probe map from expA run #2 (atlas/qwen3-0.6b-4bit/probes/v2). Intervention: layer-skip ablation (block contribution removed, residual passes through) via the Tap layer.
Hypothesis : The top-5 layers of the agreement DIFFERENTIAL map
causally support agreement behavior: skipping them
damages the model's grammatical preference
(mean logit margin of the correct verb form over
the incorrect one, on held-out minimal pairs) more
than skipping 5 random non-top layers — top-5
damage ≥ the 95th percentile of the random-5
damage distribution (20 draws), and ≥ 2× its mean.
Falsification criterion : Top-5 damage inside the random-5 distribution
(< 95th percentile) — the correlational map fails
causal verification; this survival datum (0 or 1)
is the first entry of expC's published
correlational→causal survival rate, either way.
Method : Behavioral metric: margin = mean over minimal
pairs of [logit(correct is/are) − logit(incorrect)]
at the verb position, teacher-forced prefix
"The <noun(s)> <location>". Held-out pairs: fresh
seed, deduped against every v2 probe text.
Conditions: baseline (no skip); skip top-5
differential layers; skip 5 random layers
(excluding top-5), 20 draws; skip bottom-5
differential layers (second control). Same model
(Qwen3-0.6B-4bit), deterministic forwards.
Baseline / null : baseline margin (sanity: must be > 0, i.e. the
model actually prefers grammatical agreement —
else the whole question is moot at this size);
random-5 and bottom-5 skip distributions.
Result : (pending)
Interpretation : (pending — with explicit confidence level)
Next experiment : (pending)Declared caveats. Layer-skip is a coarse intervention (removes ALL of a block's computation, not the agreement direction specifically) — a confirmed result licenses "these layers causally support agreement behavior", NOT "agreement is localized to these layers" (hydra/backup effects, notes §4.4, can mask redundancy). Direction-level interventions (LEACE-style erasure in the forward pass) are the registered follow-up.
Hypothesis — expC run #2 (specificity control + direction-level surgery)
Registered 2026-08-12 before the run, after run #1's non-confirmation. Two declared confounds get instruments: general-damage normalization for layer-skip, and a surgical direction-level intervention the probe map can legitimately pass or fail.
Hypothesis (P1) : run #1's bottom-5 damage is GENERAL, not
agreement-specific: normalizing margin damage by
general damage (mean NLL increase on neutral
prose), the bottom-5 specificity ratio falls at or
below the random-5 mean ratio, and top-5 does not
exceed the random p95 either (the run-1 verdict
stands, now with the confound measured).
Hypothesis (P2) : direction-level erasure tracks the probe map:
erasing the layer-ℓ agreement direction
(difference-in-means, estimated on agreement_A
mean-pooled reps) at all positions of ℓ's output,
minus the damage from erasing a random direction
at the same layer, yields a per-layer specific-
damage profile that correlates with the
differential probe profile: Spearman ρ ≥ 0.4
(permutation p < 0.05, 10k perms).
Falsification criterion : (P2) ρ ≤ 0 → the localization claim also dies at
direction level: survival ledger 0/2 and the
agreement map's layer structure is declared
behaviorally void at this granularity.
Method : held-out bank enlarged with 6 NEW locations
(target ≥ 180 pairs, deduped as before); NLL on
20 neutral prose sentences per condition;
direction erasure h' = h − ⟨h−μ, u⟩u applied to
every position of layer ℓ's output, u = unit
class-mean difference at ℓ, μ = grand mean;
random-direction control: 3 seeds per layer,
same procedure; all 28 layers scanned.
Baseline / null : per-layer random-direction erasure damage;
run #1 skip conditions rerun on the enlarged bank.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expC run #3 (replication of the causal direction profile)
Registered 2026-08-12 before the run. Run #2 found that erasing the diff-of-means agreement direction at any layer 2–15 destroys the behavior. By our own doctrine that profile is unpublishable until it replicates across direction estimates.
Hypothesis : The per-layer specific-damage profile is an
estimator-stable object: profiles from directions
estimated on (i) set A, (ii) set B (disjoint
nouns), (iii) 3 bootstrap resamples of set A,
agree pairwise at Spearman ρ ≥ 0.7 (mean over the
10 pairs), AND the qualitative claim holds in
EVERY source: mean specific damage over layers
2–15 ≥ 3× mean over layers 20–27.
Falsification criterion : mean pairwise ρ < 0.7 or any source violating the
3× band claim → the "causal direction profile" is
estimator noise; not publishable; expC pivots to
steering-based verification.
Method : shared random-direction controls (3 dirs/layer,
computed once); agreement-direction scan repeated
per source (5 sources × 28 layers × 192 pairs);
same held-out bank and margins as run #2.
Baseline / null : shared random-direction damage per layer.
Publication rule : if confirmed → atlas qwen3-0.6b-4bit/
interventions/v1 at Level 2 (probing evidences the
direction, erasure confirms causal load,
replicated across estimates — NOT Level 3, because
the two agreeing methods share the diff-of-means
estimator; an independent intervention family
(activation addition) is the registered L3 path).
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expC run #4 (the narrowed BAND claim)
Registered 2026-08-12 before the run, after run #3's gate refusal showed the early band replicates (+2.98…+3.34) while the late band is estimator noise. Everything fresh: direction estimates from DISJOINT HALVES of agreement_A and agreement_B (never used whole-set estimates), two new bootstrap seeds, and a NEW held-out behavioral bank (8 locations unseen by any prior run or promptset).
Hypothesis : EARLY-BAND claim only — mean specific damage over
layers 2–15 is ≥ 2.5 (on the fresh bank's
baseline-margin scale) in EVERY one of six fresh
direction sources (A-half-1, A-half-2, B-half-1,
B-half-2, bootA-fresh, bootB-fresh). The late
band (20–27) is REPORTED but carries no claim —
declared unstable by run #3.
Falsification criterion : any source with early-band mean < 2.5 → the band
claim fails replication too; the direction-erasure
program is closed for agreement at 0.6B and expC
pivots to steering-based verification.
Method : full 28-layer scan per source (map format
unchanged); shared random-direction controls
(3/layer) on the fresh bank; margins as before.
Publication rule : pass → atlas interventions/v1 published as a BAND
map at Level 2: map.json carries the six profiles,
their mean/min, the random-direction damage, and
the band statistics; the chart plots mean, worst
source, and the raw random-direction reference.
Fail → refusal documented, program closed.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expC run #5 (Level-3 path: dose-response steering)
Registered 2026-08-12 before the run. Erasure (runs #2–#4) established necessity of the agreement direction in the early band. Steering (activation addition) is the INDEPENDENT intervention family required for Level 3: if the direction is the causal handle, pushing along it must move the behavior in the predicted direction, dose-dependently, and random directions must not.
Hypothesis : At each of three early-band test layers (4, 8, 12),
adding δ = α·σℓ·u to every position of the layer
output produces a STRICTLY MONOTONE margin in
α ∈ {−2, −1, 0, +1, +2} (σℓ = std of activation
projections on u at layer ℓ), with effect size:
margin(α=−2) ≤ 0.5 × baseline at every test layer.
Specificity: the same ±2σ doses along 3 random
directions move the margin by < 25% of baseline
(mean absolute change), at every test layer.
Falsification criterion : monotonicity broken at any test layer, OR the −2σ
dose fails to halve the margin anywhere, OR random
directions move the margin ≥ 25% — steering fails,
the entry stays Level 2, and the diff-of-means
direction is declared necessary-but-not-a-handle.
Method : direction u and σℓ estimated from agreement_A
(full set — the published v1 object); behavioral
bank = run #4's fresh bank (never used for
estimation); doses applied at one layer at a time.
Publication rule : pass → atlas interventions/v2 at LEVEL 3 (erasure
necessity + steering dose-response = two
independent intervention families + probing;
v1 stays as the Level-2 record). Fail → documented,
v1 unchanged.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Hypothesis — expC run #6 (minimal L3 claim: layer 12 as a single-layer handle)
Registered 2026-08-12 before the run, after run #5's conjunction failed while L12 passed every gate. The claim is narrowed to the object that showed textbook behavior — scope discipline again, now applied to layers.
Hypothesis : At LAYER 12 ONLY, the diff-of-means agreement
direction is a dose-controlled causal handle, and
this is estimator-stable: for EVERY one of four
fresh direction sources (disjoint halves of
agreement_A and agreement_B, new permutation
seed), on a THIRD fresh behavioral bank (8
never-used locations):
(a) margins strictly monotone in
α ∈ {−2,−1,0,+1,+2}·σ;
(b) margin(−2σ) ≤ 0.5 × baseline;
(c) shared specificity control at L12: 3 random
directions at ±2σ move the margin by < 25%
of baseline on average.
Falsification criterion : any source failing (a) or (b), or (c) failing —
the single-layer L3 claim dies and Level 3 is
abandoned for this object (steering program
closed; candidate_02/expG proceed from the
Level-2 band entry).
Method : same margin metric; direction + σ re-estimated
per source at L12 from half-set mean-pooled reps;
doses applied at L12 output, all positions.
Publication rule : pass → interventions/v2 at LEVEL 3 with scope
declared as {band necessity (v1) + single-layer
L12 handle}; fail → documented refusal, v1
unchanged.
Result : (pending)
Interpretation : (pending)
Next experiment : (pending)Analysis
Analysis — expC run #1: the agreement probe map FAILS causal verification
Run: results/expC_causal_verification/20260812T064534Z/results.json (5.7 s).
Model: Qwen3-0.6B-4bit. Input map: agreement differential (probes/v2).
Intervention: layer-skip ablation via the Tap layer. Hypothesis and the
binary verdict criterion registered before the run.
Hypothesis : the top-5 differential layers (by real−twin probe
selectivity) causally support agreement behavior —
skip damage ≥ p95 of random-5 draws AND ≥ 2× their
mean.
Falsification criterion : top-5 damage inside the random distribution.
Result : FALSIFIED — survival 0/1. Baseline grammatical
margin +4.63 (sanity holds: the model robustly
prefers correct agreement). Skip damage:
top-5 (layers 17,18,19,21,22) = +2.17 —
BELOW the random-5 mean (+3.14, p95 +4.85, 20
draws); bottom-5 differential (layers 0–4) = +4.99,
the LARGEST of all conditions.
Interpretation : Level 1 for the negative claim. Where agreement
information is most linearly decodable above the
architecture null (late-mid layers) is NOT where
the computation is causally load-bearing for the
behavior. This is the Hase-class dissociation
(localization ≠ causal support — notes §4.6)
measured end-to-end in our own pipeline, on a
pre-registered binary verdict. The
correlational→causal survival ledger opens at 0/1.
Declared caveats bite exactly as registered:
(a) layer-skip is coarse — early-layer skips
plausibly cause GENERAL degradation, not
agreement-specific damage (bottom-5 +4.99 reads as
"the model breaks", not "agreement lives at layers
0–4"); a specificity control (margin damage
normalized by general perplexity damage) is the
registered follow-up; (b) held-out pairs = 89
after dedup against every v2 probe text (below the
planned 200 — combo pool exhausted; enlarge banks
next run).
Next experiment : run #2 with (1) perplexity-normalized specificity
scores per skip condition, (2) direction-level
intervention (LEACE erasure of the agreement
direction in the forward pass at layer ℓ) — a
surgical test the probe map CAN legitimately pass
or fail, (3) larger held-out bank. Then the same
protocol on arith_valid.Ledger
| # | Correlational claim | Intervention | Survives? |
|---|---|---|---|
| 1 | agreement top-5 differential layers (probes/v2) | layer-skip ablation | No (0/1) |
| 2 | agreement differential layer profile (probes/v2) | direction-level erasure scan | No (0/2) |
The published survival rate is the running fraction of this table — the charter §8.3 metric, now live.
Analysis — expC run #2: ledger 0/2 — and the control scan finds the real structure
Run: results/expC_causal_verification/20260812T065406Z/results.json (192
held-out pairs, baseline margin +4.24, NLL 5.73). Hypotheses registered
before the run.
Hypothesis (P1) : run #1's bottom-5 damage is general, not
agreement-specific, once normalized by NLL damage.
Result (P1) : CONFIRMED. Specificity (margin damage / NLL
damage): bottom-5 = 0.51 — BELOW the random-5 mean
(1.44); top-5 = 2.51 — above the random mean but
below the p95 (4.73). Run #1's verdict stands, now
with the confound measured: early-layer skips
break the model generally; nothing in the skip
family singles out the probe map's top layers.
Hypothesis (P2) : per-layer direction-erasure specific damage
correlates with the differential probe profile
(Spearman ρ ≥ 0.4, perm-p < 0.05).
Result (P2) : FALSIFIED, decisively — ρ = −0.136, perm-p 0.757.
Survival ledger: 0/2. The probe map's layer
ranking is behaviorally void at this granularity.
BUT the scan itself uncovered strong causal
structure the probe map missed: erasing the
layer-ℓ agreement direction (diff-of-means, with
random-direction controls netted out) at ANY
single layer in 2–15 destroys most of the margin
(specific damage +2.6…+4.0 of a +4.24 baseline;
L12: +3.97, L15: +3.75, L6: +3.71), while the
late layers the probes ranked highest carry
little (L21: +0.69, L20: +0.29) — and L18/L22
erasure slightly HELPS (−1.86/−1.09), suggesting
suppressive components.
Interpretation : Level 0–1 (single model, single direction
estimate from set A, no seed replication of the
intervention yet — by our own doctrine this
causal profile is NOT publishable as an atlas
entry until replicated). Two lessons:
(1) correlational layer rankings did not survive
two different causal tests — the survival rate
the field never publishes is, so far, 0%;
(2) the behaviorally load-bearing object is a
low-dimensional DIRECTION present across
early-mid layers, not a "place" — consistent with
the linear-representation view and with why
decodability peaks (late, after information is
everywhere) diverge from causal joints (early,
where the direction is constructed). Random-
direction erasure also hurts at layers 2–9
(+0.9…+2.5): early residual streams are fragile
to ANY rank-1 deletion — netted out in the
specific column.
Next experiment : run #3 — replicate the direction-erasure profile
(direction re-estimated on set B and on 5
bootstrap seeds; report profile replication rate)
→ if stable, publish as the atlas's first
INTERVENTIONS map (Level toward 2–3) and make the
probes/v2 entry point to it as the causal
counterpart. Then the same scan on arith_valid.Analysis — expC run #3: the publication gate REFUSED the causal profile
Run: results/expC_causal_verification/20260812T065929Z/results.json.
Five direction sources (set A, set B, 3 bootstraps of A), shared
random-direction controls, same held-out bank (192 pairs, baseline +4.24).
Hypothesis, criterion, and publication rule registered before the run.
Hypothesis : the per-layer specific-damage profile is
estimator-stable — mean pairwise Spearman ρ ≥ 0.7
and early(2–15) ≥ 3× late(20–27) in EVERY source.
Result : FALSIFIED — mean ρ = 0.495 (min 0.176); the band
claim fails for setB and bootA2 (ratio 1.2).
make_interventions_mapcard.py refused the entry
(exit 1), per the registered rule. The atlas
stays at two entries — the gate did its job.
THE STABLE PART: early-band (layers 2–15) mean
specific damage replicates tightly across all
five sources (+2.98, +3.22, +3.34, +3.27, +3.28)
— erasing the estimated agreement direction
anywhere in the early-mid band reliably destroys
~70–79% of the behavior, whatever the estimation
set. THE UNSTABLE PART: the late band (20–27)
swings from −0.44 to +2.77 with the estimator —
the late-layer "profile" is direction-estimation
noise, which also retroactively explains run #2's
anti-correlation (the probe ranking lives exactly
where the causal profile is noise).
Interpretation : Level 0–1. The full profile is NOT a stable
object; the coarser claim ("an early-band
agreement direction is causally load-bearing") is
the replicable candidate. Publishing gates that
refuse are the mechanism that keeps the atlas
honest — this refusal is itself a process result
worth reporting on the site's methodology page.
Next experiment : run #4 with the NARROWER pre-registered claim:
early-band (2–15) mean specific damage ≥ 2.5 on
fresh direction estimates (new bootstrap seeds +
a held-out estimation split) and a fresh
behavioral bank; late band explicitly excluded as
unstable. If it passes, publish the interventions
map as an early-band BAND claim (not a per-layer
ranking), Level 2.Analysis — expC run #4: the BAND claim passes — first Level-2 atlas entry
Run: results/expC_causal_verification/20260812T071954Z/results.json.
Everything fresh: six direction sources (disjoint halves of A and B + two
new bootstraps), new behavioral bank (8 unseen locations, 192 pairs,
baseline margin +4.45). Claim, bar, and publication rule registered before
the run.
Hypothesis : early-band (2–15) mean specific damage ≥ 2.5 in
EVERY fresh source; late band reported, no claim.
Result : CONFIRMED — early-band means: Ahalf1 +3.35,
Ahalf2 +3.23, bootA +3.30, Bhalf1 +3.26,
Bhalf2 +3.33, bootB +3.27 (min +3.228 vs bar
2.5). Spread across six estimators: 0.12 — the
band is a tight, estimator-stable causal object.
Late band again unstable (+0.27…+2.83), as
declared; it carries no claim.
Interpretation : Level 2, published: atlas/qwen3-0.6b-4bit/
interventions/v1 — the project's first Level-2
entry and first causal map. The claim: a single
linear direction (diff-of-means over correct vs
violated agreement), erased at ANY one layer in
2–15, removes ~73–75% of the model's grammatical
preference on held-out pairs, replicated across
six independent direction estimates and netted
against random-direction damage. NOT Level 3:
both agreeing methods share the diff-of-means
estimator — activation-addition steering is the
registered L3 path. Granularity discipline paid
off: the per-layer version of this map was
refused (run #3); the band version replicates.
Next experiment : (1) L3 path: steering (add the direction) should
INCREASE the margin on violated-preference pairs;
(2) same band protocol on arith_valid; (3)
quantization drift of the band map (candidate_02:
does the early band move under Q8/FP16?);
(4) cross-model: does the band replicate on
Qwen3-1.7B (expG entry)?Analysis — expC run #5: Level-3 gate refused — and layer 12 is a textbook handle
Run: results/expC_causal_verification/20260812T072455Z/results.json.
Steering doses α·σℓ·u at layers {4, 8, 12}, α ∈ {−2…+2}; conjunctive
criterion (monotone AND halved AND specific at EVERY test layer) registered
before the run. make_l3_mapcard.py refused (exit 1); v1 stays Level 2.
Hypothesis : strict dose-response monotonicity + halving +
random-direction specificity at all of L4/L8/L12.
Result : FALSIFIED as a conjunction.
L12 — PASSES EVERYTHING, textbook: margins
+1.95 / +3.00 / +4.45 / +5.82 / +6.79 across
−2σ…+2σ (strictly monotone), halved at −2σ,
random-direction |Δ| = 0.45 vs bound 1.11.
L08 — near-monotone (+2σ dips: 5.70 < 6.26),
random |Δ| 2.97: not specific at this dose.
L04 — OVERDOSE REGIME: ±2σ both collapse the
margin (+0.08 / +0.04) and random directions are
equally destructive (|Δ| 4.40): at early layers,
ANY perturbation of magnitude ~2σ (σ estimated
from mean-pooled reps) breaks the computation —
the rank-1 fragility of run #2, now dose-resolved.
Interpretation : Level 0–1. The direction is a clean, dose-
controlled causal handle at mid-band (L12) but
the conjunctive claim over the whole test set
fails, so the gate held v2 back — correctly.
Reading: "necessity everywhere in the band"
(erasure, L2 entry) coexists with "controllable
handle only where the layer tolerates
perturbation". Dose scale is a confound at early
layers: σ from mean-pooled statistics likely
overdoses positions at layers with different
norm profiles.
Next experiment : run #6 (registered idea): layer-local dose
calibration (σ from position-level projections at
the target layer; or a dose-sweep to find each
layer's non-destructive range), then re-register
the handle claim on the sub-band that tolerates
calibrated doses (candidate: 10–14). Also register
the L12 single-layer handle claim with fresh
direction estimates as a minimal L3 candidate.Gate record (for methodology.md)
Three refusals/passes to date, all pre-registered: run #3 per-layer profile REFUSED → run #4 band claim PASSED (Level 2 published) → run #5 L3 conjunction REFUSED (v1 unchanged). The atlas never received a claim its evidence didn't carry.
Analysis — expC run #6: the L12 handle claim fails replication — Level 3 abandoned
Run: results/expC_causal_verification/20260812T072913Z/results.json.
Four fresh direction sources at L12, third fresh bank (192 pairs, baseline
+4.57), fresh random-direction seeds. Registered rule: any source failing
monotonicity/halving, or specificity failing, kills the L3 claim and closes
the steering program for this object.
Hypothesis : L12 is an estimator-stable dose-controlled handle.
Result : FALSIFIED. Monotone: 1/4 sources only (Ahalf2);
the POSITIVE dose arm is unstable (Ahalf1 +2σ dips
5.59<6.29; Bhalf1 +1σ dips below baseline; Bhalf2
+2σ < +1σ). Halving at −2σ: 4/4 — the negative
(erasure-like) arm is robust, again. SPECIFICITY
FAILED TO REPLICATE: random-direction mean |Δ| =
3.36 vs bound 1.14 on the fresh bank and fresh
random seeds (run #5's L12 value was 0.45 — with
only 3 random draws, that pass now reads as
sampling luck).
Interpretation : Level 1 for the negative. The direction is
NECESSARY (v1, Level 2, band-replicated) but NOT
a reliable additive handle: pushing along it does
not control the behavior in a dose-stable,
direction-specific way. Per the pre-registered
rule, Level 3 is abandoned for this object and
the steering program is closed. Had run #5's L12
observation been published without fresh
re-registration, the atlas would now contain a
false Level-3 claim — the gate earned its keep a
third time.
Methodology lesson : 3 random-direction draws are too few for a
specificity bound; methodology.md will require
≥10 draws with a percentile bound for any
specificity control from now on.
Next experiment : program closed here. Proceeding tracks: arith_valid
band protocol; candidate_02 (quantization drift of
the Level-2 band map); expG (band replication on
Qwen3-1.7B). The final gate record for this arc:
refuse → pass(L2) → refuse → refuse.README
expC_causal_verification
Causal verification pipeline — correlational-to-causal survival rate
Status: scaffolded — hypothesis to be registered before any benchmark runs (charter §10).
Result runs
- 20260812T072913Z / results.json 2.5 KiB
- 20260812T072455Z / results.json 2.5 KiB
- 20260812T071954Z / results.json 8.4 KiB
- 20260812T065929Z / results.json 5.8 KiB
- 20260812T065406Z / results.json 11.6 KiB
- 20260812T064534Z / results.json 2.0 KiB