Results / atlas/qwen3-0.6b-4bit/probes/v2/confidence.md

---
project: modelmap
document: qwen3-0.6b-4bit/probes/v2 — confidence
author: Simon-Pierre Boucher
contact: contact@spboucher.ai
website: https://modelmap.io
created: 2026-08-12
status: reviewed
---

# Confidence — qwen3-0.6b-4bit / probes / v2

```text
Level      : 1
Seeds      : 5
Prompt sets: 6 (token-balanced, structure-borne; overlap certificates in manifest)
Methods in agreement : 1 (linear probes only — Level 2 requires a second method)
Causal verification  : ATTEMPTED AND FAILED (expC run #1 layer-skip: top-5
                       differential layers not confirmed; survival 0/1)
```

Per-property verdicts (differential real−twin, mean pooling):
- word_order: null-dominated (twin acc 0.96; surface statistics explain the map); maxAcc A (mean-pool) 0.996, twin acc 0.958, signal layers 2/28
- agreement: trained-model signal (real−twin sel > 0.10 on 25/28 layers, max +0.38); maxAcc A (mean-pool) 0.967, twin acc 0.729, signal layers 25/28
- arith_valid: trained-model signal (real acc 0.86–0.90 vs twin 0.56–0.58); maxAcc A (mean-pool) 0.858, twin acc 0.558, signal layers 5/28

Published claims are DIFFERENTIAL only (real minus random-init twin), per the
doctrine adopted after v1's validity-gate failure. The strict twin gate
(selectivity < 0.05) still fails on word_order and agreement — the twin
extracts real surface signal from tokenization statistics — so raw probe
accuracies are never cited as evidence of learned structure. What survives:
agreement and arith_valid show layer-resolved trained-model signal
(Level 1, correlational; 5 seeds × 2 sets, controls listed above).