Atlas / atlas/qwen3-0.6b-4bit/probes/v2
qwen3-0.6b-4bit — probes v2
Level 1 — correlational
The map
trained · set Atrained · set Brandom-init twin
trained · set Atrained · set Brandom-init twin
trained · set Atrained · set Brandom-init twin
Confidence
Confidence — qwen3-0.6b-4bit / probes / v2
Level : 1
Seeds : 5
Prompt sets: 6 (token-balanced, structure-borne; overlap certificates in manifest)
Methods in agreement : 1 (linear probes only — Level 2 requires a second method)
Causal verification : ATTEMPTED AND FAILED (expC run #1 layer-skip: top-5
differential layers not confirmed; survival 0/1)Per-property verdicts (differential real−twin, mean pooling):
- word_order: null-dominated (twin acc 0.96; surface statistics explain the map); maxAcc A (mean-pool) 0.996, twin acc 0.958, signal layers 2/28
- agreement: trained-model signal (real−twin sel > 0.10 on 25/28 layers, max +0.38); maxAcc A (mean-pool) 0.967, twin acc 0.729, signal layers 25/28
- arith_valid: trained-model signal (real acc 0.86–0.90 vs twin 0.56–0.58); maxAcc A (mean-pool) 0.858, twin acc 0.558, signal layers 5/28
Published claims are DIFFERENTIAL only (real minus random-init twin), per the doctrine adopted after v1's validity-gate failure. The strict twin gate (selectivity < 0.05) still fails on word_order and agreement — the twin extracts real surface signal from tokenization statistics — so raw probe accuracies are never cited as evidence of learned structure. What survives: agreement and arith_valid show layer-resolved trained-model signal (Level 1, correlational; 5 seeds × 2 sets, controls listed above).
Provenance
{
"author": "Simon-Pierre Boucher",
"contact": "contact@spboucher.ai",
"website": "https://modelmap.io",
"model_id": "mlx-community/Qwen3-0.6B-4bit",
"map_type": "probes",
"version": "v2",
"commit": "1dcd820126ca96fe99c1b473947f3c7e54038500",
"model_hash": "392e8d466d56100ada00eb82031fb854297fc9e389b7d303eba3af114e87bce2",
"config": {
"model": "mlx-community/Qwen3-0.6B-4bit",
"seeds": [
0,
1,
2,
3,
4
],
"top_k": 5,
"fdr_q": 0.05,
"poolings": [
"mean",
"last"
],
"promptsets": {
"version": "v2",
"seed": 54321,
"n_per_class": 120,
"author": "Simon-Pierre Boucher",
"contact": "contact@spboucher.ai",
"website": "https://modelmap.io",
"design": "structure-borne properties, token-balanced classes (response to expA run #1 validity-gate failure)",
"files": {
"agreement_A.jsonl": {
"sha256": "622c5e0966d2b8e8abc9fb7b964a8e5de8df5f8d41c37a0ed4ed8c8894f1f0ea",
"n": 240,
"class_token_overlap": 1
},
"agreement_B.jsonl": {
"sha256": "243b85e2c207be6cff860a4660400c9c90859d5014636177b4475959b872542f",
"n": 240,
"class_token_overlap": 1
},
"arith_valid_A.jsonl": {
"sha256": "a4bb4c645b3f1797e562f62765d04e7603d0e0ac0286e013cfa90f774de27159",
"n": 240,
"class_token_overlap": 0.6735
},
"arith_valid_B.jsonl": {
"sha256": "adc68ef876a823ec43c4cf3c64a408f46b2613e2f3e7386fa1d20e3f80e9cf93",
"n": 240,
"class_token_overlap": 0.9674
},
"word_order_A.jsonl": {
"sha256": "9052b924930aa5cde7ce434439c418cbd3e5a0731d421cd370bd46d60abcdb66",
"n": 240,
"class_token_overlap": 0.5294
},
"word_order_B.jsonl": {
"sha256": "3012db3b5053ba5d113d0fe6529c4b88b361b9e4572044b9e5e64a3480929a1f",
"n": 240,
"class_token_overlap": 0.5072
}
}
}
},
"seed": [
0,
1,
2,
3,
4
],
"hardware_manifest": {
"author": "Simon-Pierre Boucher",
"contact": "contact@spboucher.ai",
"website": "https://modelmap.io",
"chip": {
"brand": "Apple M5 Max",
"cores_total": 18,
"cores_performance": 6,
"cores_efficiency": 12
},
"memory": {
"unified_gb": 48,
"pagesize": 16384
},
"os": {
"system": "Darwin",
"version": "27.0",
"arch": "arm64"
},
"software": {
"python": "3.14.4",
"numpy": "2.5.2",
"mlx": "0.32.0",
"torch": "2.13.0",
"safetensors": "0.8.0"
}
},
"created": "2026-08-12",
"source_results": "results/expA_probe_reliability/20260812T063856Z/results.json"
}Files
- confidence.md 1.6 KiB
- map.json 171.0 KiB
- mapcard.json 3.0 KiB
- provenance.json 2.8 KiB