
We present XRayBench v0.2, a seven-phase controlled benchmark for evaluating option-scoring evidence about mathematical representations in transformer language models. The executed training battery contains 50 tasks in 5 mathematical families, then evaluates each signal against three control conditions (random-label, matched-token, semantics-breaking), paraphrase stability, activation patching, and MLP/attention ablation; ablation phases are implemented but their results are not reported in v0.2. We define the XRay Score, a single metric in [-4, +4] summarizing reported evidence phases, and introduce the Partial Structure Index (PSI), a per-task metric for residual option-scoring signal beyond the random-label and matched-token controls. For distilgpt2 (6 blocks, 82M), GPT-2 (12 blocks, 117M), and GPT-2-medium (24 blocks, 355M), XRay Scores are non-positive. Confound dominance varies by checkpoint, with matched-token controls producing the largest gap for GPT-2. PSI-only analyses additionally cover GPT-2-large (36 blocks, 774M) and Pythia models from 70M to 1B parameters. In these measurements, PSI shows a family-dependent pattern: the unweighted mean over binder\_tracking, theorem\_pattern, and proof\_closure is higher, but non-monotone, than the unweighted mean over arithmetic and algebra at several tested scales. A descriptive threshold rule reaches 98\% accuracy using n_real_win_layers, a direct PSI component. The pattern is examined on held-out parameterizations and preliminarily cross-checked on Pythia under a different, first-token protocol. The global XRay Score masks a family-dependent pattern in the tested measurements: logic and computation show different PSI profiles over the evaluated models. Keywords: mechanistic interpretability, mathematical reasoning, controlled benchmark, logit lens, token bias, representation probe, partial structure index, emergent representations Maturity: Draft. Target venue: NeurIPS 2026 (Datasets and Benchmarks Track). Part of The Latent research program.
