Results / atlas/qwen3-0.6b-4bit/probes/v1/confidence.md

---
project: modelmap
document: qwen3-0.6b-4bit/probes/v1 — confidence
author: Simon-Pierre Boucher
contact: contact@spboucher.ai
website: https://modelmap.io
created: 2026-08-12
status: reviewed
---

# Confidence — qwen3-0.6b-4bit / probes / v1

```text
Level      : 1
Seeds      : 5
Prompt sets: 6 (2 disjoint template families per property)
Methods in agreement : 1 (linear probes only — Level 2 requires a second method)
Causal verification  : none (observational; Level 3 requires intervention)
```

Per-property evidence:
- lang_id: maxAcc A/B = 1.000/1.000, seed SD 0.0000, dataset shift 0.0000, twin max selectivity 0.808, replication(top-5) 1.00/1.00
- code_prose: maxAcc A/B = 1.000/1.000, seed SD 0.0000, dataset shift 0.0000, twin max selectivity 0.883, replication(top-5) 1.00/1.00
- arith: maxAcc A/B = 1.000/1.000, seed SD 0.0000, dataset shift 0.0000, twin max selectivity 0.558, replication(top-5) 1.00/1.00

**This is a published NEGATIVE result (Level 1 for the negative claim).**
The random-init architecture twin matches the trained model at ceiling
(accuracy 1.00, 28/28 layers FDR-significant, for the twin as for the real
model; mean real-minus-twin selectivity within +/-0.06). By the validity
criterion registered in hypothesis.md BEFORE the run (twin selectivity must
stay < 0.05), this probing harness is INVALID for localization claims on
these promptsets: it measures the tokenizer + architecture prior, not
learned computation. The negative claim itself is controlled and replicated
(5 seeds, 2 disjoint promptsets, 3 properties) - hence Level 1.

Consequences adopted: (1) probe maps are only publishable as REAL-MINUS-TWIN
differentials; (2) promptsets v2 must remove lexical separability (shared
vocabulary across classes); (3) the seed-vs-dataset variance hypothesis is
untestable at ceiling and moves to run #2.