
Across 36 (technique × model) cells at n=400 per cell (detecting ≥5 pp effects at 80% power), no cell exceeded plain+CoT with BH-significant support (a single cell, Qwen3-8B + Plan-and-Solve, shows a raw paired Δ of +2.25 pp that does not survive BH at k=36: p=.150, qBH=.504, and falls within the audit's undetectable-effect band); three cells on Phi-4 significantly underperformed after Benjamini–Hochberg correction at q<0.05. A mechanistic decomposition attributed the full apparent benefit to CoT alone (+2.23 pp), with persona scaffolding and three-sample Self-Consistency contributing null effects. Eight inter-model aggregation strategies (four in-pool-verifier, four with out-of-pool GPT-4o verifier) all failed on held-out data, with deltas from -0.38 pp to -9.50 pp. Qwen3-30B-A3B with plain CoT achieves 79.98% on the full 12,032-question MMLU-Pro benchmark, exceeding GPT-4o's published five-shot-CoT leaderboard value (72.60%) by 7.38 pp at approximately 25× lower inference cost (OpenRouter list pricing), and tying DeepSeek-V3 (80.46%) within noise at comparable pricing. The GPT-4o comparison is protocol-asymmetric and not the primary headline. A contamination-matched frontier peer (Claude Sonnet 4.6) led Qwen3-30B-A3B by +6.42 pp (McNemar p=2.08×10-10) on a 3-seed stratified 1,200-item paired sample, replicating the within-model plain-CoT gain at frontier scale. We adopt the label Compute Confound for this specific compound-prompting instance of the compute-asymmetric evaluation problem. Deployment recommendation (for MCQ reasoning at 8B–30B small-open-source scale): a single best-in-class small model plus plain CoT. Within the audited wrapper family on MCQ reasoning at this scale, stop adding wrappers unless compute-matched lift on held-out data can be demonstrated under Benjamini–Hochberg correction. Project page: https://bgml.ai Source code: https://github.com/BGMLAI
small language models, Qwen3, compute-matched evaluation, compound prompting, chain-of-thought reasoning, large language models, Benjamini-Hochberg correction, Phi-4, McNemar test, benchmark auditing, consumer hardware, open-source LLM, MMLU-Pro
small language models, Qwen3, compute-matched evaluation, compound prompting, chain-of-thought reasoning, large language models, Benjamini-Hochberg correction, Phi-4, McNemar test, benchmark auditing, consumer hardware, open-source LLM, MMLU-Pro
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
