
We evaluate Q-Compass 1 β a value-projection-free attention primitive grounded in the reinforcement-learning π-function β at three parameter scales (57M, 121M, 179M) across four modalities (WikiText-103, MS-COCO captions, LibriSpeech clean-100, MiniGrid 3D navigation). We evaluate the SAVO four-projection variant in which the π projects the stateβaction product instead of the raw input. We also evaluate multi-head Q-Compass (MH-QC) and report it as a null result. Text-LM parity (controlled ablation): SAVO sits +12.33 Β± 0.87 perplexity above the rank-matched transformer at 60m (paired-difference, 4 seeds, π = 7.6 Γ 10β4 ). Full-rank standard 8-head MHA at the same training recipe (2 seeds) reaches 257.96 Β± 2.12 val ppl β worse than the rank-matched controlled ablation at this 10,000-step budget, despite having 8Γ more attention-block parameters. SAVO is +5.79 ppl above full-rank MHA on val. The same βΌ12-ppl gap to rank-matched holds at 120m and 180m (single seed each). Cross-modal non-interference: at matched 229M-text-token compute, joint four-modality 60m training reaches the same per-text-token loss as text-only training, and the property holds through 180m. Out-of-distribution: the 60m ππ -free SAO has a small OOD edge on arxiv and pubmed; the effect does not replicate at 120m or 180m. We report this as a null result on the ππ -free OOD-generalisation hypothesis at the scales tested. Cross-field demonstration: the same SAVO block class runs on four computational-oncology tasks β signature decomposition (cosine 0.975 vs NNLS 0.987, NNLS higher), 27-class pan-cancer (top-1 0.517 vs majority 0.087), GDSC2 drug-response (Pearson 0.903; drug-only baseline 0.864), and TCGA 5-year-survival (Cindex 0.701; clinical-only ablation 0.708, clinical-only higher within seed noise). World-model branch: the world objective trains concurrently with text/vision/audio without breaking the joint training; world MSE drops from 1.125 at initialisation to 0.071 at step 10,000 (βΌ16Γ reduction). The predict-mean baseline on the trained 180m StateEncoder is 0.033 in the same encoded-state metric, reflecting MiniGrid-Empty-8x8βs low next-state variance β benchmarking the routing primitive against dedicated world-model architectures on demanding environments (DMLab, Habitat) is out of scope. The unification claim is a structural property of the routing block. NanoG1 (cancer foundation model with mid-CoT hypothetical simulation, building on Β§7) is deferred to a subsequent paper
reinforcement learning, Q-function, language model, transformer, sequence mixing, World Modeling, attention mechanism
reinforcement learning, Q-function, language model, transformer, sequence mixing, World Modeling, attention mechanism
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
