Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

Compound Prompting Vanishes Under Matched Compute: A Twelve-Wrapper Audit of Small Language Models on MMLU-Pro

Authors: Jankiewicz, Bogumił;

Compound Prompting Vanishes Under Matched Compute: A Twelve-Wrapper Audit of Small Language Models on MMLU-Pro

Abstract

Across 36 (technique × model) cells at n=400 per cell (detecting ≥5 pp effects at 80% power), no cell exceeded plain+CoT with BH-significant support (a single cell, Qwen3-8B + Plan-and-Solve, shows a raw paired Δ of +2.25 pp that does not survive BH at k=36: p=.150, qBH=.504, and falls within the audit's undetectable-effect band); three cells on Phi-4 significantly underperformed after Benjamini–Hochberg correction at q<0.05. A mechanistic decomposition attributed the full apparent benefit to CoT alone (+2.23 pp), with persona scaffolding and three-sample Self-Consistency contributing null effects. Eight inter-model aggregation strategies (four in-pool-verifier, four with out-of-pool GPT-4o verifier) all failed on held-out data, with deltas from -0.38 pp to -9.50 pp. Qwen3-30B-A3B with plain CoT achieves 79.98% on the full 12,032-question MMLU-Pro benchmark, exceeding GPT-4o's published five-shot-CoT leaderboard value (72.60%) by 7.38 pp at approximately 25× lower inference cost (OpenRouter list pricing), and tying DeepSeek-V3 (80.46%) within noise at comparable pricing. The GPT-4o comparison is protocol-asymmetric and not the primary headline. A contamination-matched frontier peer (Claude Sonnet 4.6) led Qwen3-30B-A3B by +6.42 pp (McNemar p=2.08×10-10) on a 3-seed stratified 1,200-item paired sample, replicating the within-model plain-CoT gain at frontier scale. We adopt the label Compute Confound for this specific compound-prompting instance of the compute-asymmetric evaluation problem. Deployment recommendation (for MCQ reasoning at 8B–30B small-open-source scale): a single best-in-class small model plus plain CoT. Within the audited wrapper family on MCQ reasoning at this scale, stop adding wrappers unless compute-matched lift on held-out data can be demonstrated under Benjamini–Hochberg correction. Project page: https://bgml.ai Source code: https://github.com/BGMLAI

Keywords

small language models, Qwen3, compute-matched evaluation, compound prompting, chain-of-thought reasoning, large language models, Benjamini-Hochberg correction, Phi-4, McNemar test, benchmark auditing, consumer hardware, open-source LLM, MMLU-Pro

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green