Atlas / atlas/qwen3-0.6b-4bit/probes/v1
qwen3-0.6b-4bit — probes v1
Level 1 — correlational
The map
Confidence
Confidence — qwen3-0.6b-4bit / probes / v1
Level : 1
Seeds : 5
Prompt sets: 6 (2 disjoint template families per property)
Methods in agreement : 1 (linear probes only — Level 2 requires a second method)
Causal verification : none (observational; Level 3 requires intervention)Per-property evidence:
- lang_id: maxAcc A/B = 1.000/1.000, seed SD 0.0000, dataset shift 0.0000, twin max selectivity 0.808, replication(top-5) 1.00/1.00
- code_prose: maxAcc A/B = 1.000/1.000, seed SD 0.0000, dataset shift 0.0000, twin max selectivity 0.883, replication(top-5) 1.00/1.00
- arith: maxAcc A/B = 1.000/1.000, seed SD 0.0000, dataset shift 0.0000, twin max selectivity 0.558, replication(top-5) 1.00/1.00
This is a published NEGATIVE result (Level 1 for the negative claim). The random-init architecture twin matches the trained model at ceiling (accuracy 1.00, 28/28 layers FDR-significant, for the twin as for the real model; mean real-minus-twin selectivity within +/-0.06). By the validity criterion registered in hypothesis.md BEFORE the run (twin selectivity must stay < 0.05), this probing harness is INVALID for localization claims on these promptsets: it measures the tokenizer + architecture prior, not learned computation. The negative claim itself is controlled and replicated (5 seeds, 2 disjoint promptsets, 3 properties) - hence Level 1.
Consequences adopted: (1) probe maps are only publishable as REAL-MINUS-TWIN differentials; (2) promptsets v2 must remove lexical separability (shared vocabulary across classes); (3) the seed-vs-dataset variance hypothesis is untestable at ceiling and moves to run #2.
Provenance
{
"author": "Simon-Pierre Boucher",
"contact": "contact@spboucher.ai",
"website": "https://modelmap.io",
"model_id": "mlx-community/Qwen3-0.6B-4bit",
"map_type": "probes",
"version": "v1",
"commit": "3935e7933294b9c1293cd31b887126402fc53115",
"model_hash": "392e8d466d56100ada00eb82031fb854297fc9e389b7d303eba3af114e87bce2",
"config": {
"model": "mlx-community/Qwen3-0.6B-4bit",
"seeds": [
0,
1,
2,
3,
4
],
"top_k": 5,
"fdr_q": 0.05,
"promptsets": {
"version": "v1",
"seed": 12345,
"n_per_class": 120,
"author": "Simon-Pierre Boucher",
"contact": "contact@spboucher.ai",
"website": "https://modelmap.io",
"limitation": "template-generated v1; natural-corpus v2 registered",
"files": {
"arith_A.jsonl": {
"sha256": "2fd80600d8a1b4ad89fed0cdeba2d5d6c33e795f7000553c08addc361305e074",
"n": 240
},
"arith_B.jsonl": {
"sha256": "04bb0eceb4b9262e090cd45b11d77487e438d14859a809d98377ac1807e3a502",
"n": 240
},
"code_prose_A.jsonl": {
"sha256": "dfbfb13dade0fd0be1202b0a745cfd03dfa796f2de4e22c6eca116abab71415c",
"n": 240
},
"code_prose_B.jsonl": {
"sha256": "e9c3f78d8754b0ad5c8921ed361fd4a4a8e824a3c116d2cb3f372417793fd4d7",
"n": 240
},
"lang_id_A.jsonl": {
"sha256": "43bd7ed12d2d9f95d643556a1a05a75623f32201b6f47fdc7d580b5f0cdce30a",
"n": 240
},
"lang_id_B.jsonl": {
"sha256": "833435dae9a61ae04cf354484693c821f6a88dfceaf19865b9da2947c9d71b52",
"n": 240
}
}
}
},
"seed": [
0,
1,
2,
3,
4
],
"hardware_manifest": {
"author": "Simon-Pierre Boucher",
"contact": "contact@spboucher.ai",
"website": "https://modelmap.io",
"chip": {
"brand": "Apple M5 Max",
"cores_total": 18,
"cores_performance": 6,
"cores_efficiency": 12
},
"memory": {
"unified_gb": 48,
"pagesize": 16384
},
"os": {
"system": "Darwin",
"version": "27.0",
"arch": "arm64"
},
"software": {
"python": "3.14.4",
"numpy": "2.5.2",
"mlx": "0.32.0",
"torch": "2.13.0",
"safetensors": "0.8.0"
}
},
"created": "2026-08-12",
"source_results": "results/expA_probe_reliability/20260812T062605Z/results.json"
}Files
- confidence.md 1.8 KiB
- map.json 63.5 KiB
- mapcard.json 2.5 KiB
- provenance.json 2.4 KiB