Results / atlas/qwen3-0.6b-4bit/interventions/v1/confidence.md
---
project: modelmap
document: qwen3-0.6b-4bit/interventions/v1 — confidence
author: Simon-Pierre Boucher
contact: contact@spboucher.ai
website: https://modelmap.io
created: 2026-08-12
status: reviewed
---
# Confidence — qwen3-0.6b-4bit / interventions / v1
```text
Level : 2
Seeds : six fresh direction sources (disjoint halves of two promptsets
+ two fresh bootstraps); fresh behavioral bank
Prompt sets: 3 (two estimation sets + held-out behavioral bank)
Methods in agreement : 2 (diff-of-means probing; direction erasure) — shared
estimator, hence Level 2 and not 3
Causal verification : YES — rank-1 erasure with random-direction nulls
```
Pre-registered band claim (bar 2.5 on a +4.45 baseline margin):
- Ahalf1: early-band mean +3.347
- Ahalf2: early-band mean +3.228
- bootA: early-band mean +3.302
- Bhalf1: early-band mean +3.263
- Bhalf2: early-band mean +3.326
- bootB: early-band mean +3.270
Minimum early-band mean across sources: +3.228 — claim PASSES.
Granularity discipline: run #3's per-layer profile FAILED replication and was
refused by this very gate; the published object is the BAND (layers 2–15).
The late band (20–27) is displayed but carries no claim. This map is the
causal counterpart of probes/v2, whose decodability ranking it contradicts
(survival ledger 0/2) — both stay published, labeled by what they measure.