Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
addClaim

You Are Wasting 96% of Your Inference Compute: 23x Intelligence Per Watt from Software Alone — A Multiplicative Algorithmic Stack for LLM Inference

Authors: Cantrell, Cole;

You Are Wasting 96% of Your Inference Compute: 23x Intelligence Per Watt from Software Alone — A Multiplicative Algorithmic Stack for LLM Inference

Abstract

Saad-Falcon et al. (2025) introduced Intelligence Per Watt (IPW) as the critical metric for tracking AI efficiency: task accuracy divided by power consumed. Their longitudinal study documents 5.3x IPW improvement from 2023-2025, driven by model and hardware advances. This paper demonstrates that a stack of algorithmic efficiency optimizations — derived from a unified stochastic health monitoring framework — provides an additional multiplicative IPW improvement on top of whatever hardware is available. The core three-layer algorithmic stack (FlashAttention, run-level power metric allocation, and early exit) provides a combined 23x IPW improvement at sequence length 4,096 tokens, through three orthogonal mechanisms: per-operation memory bandwidth efficiency (FlashAttention, 2.86x), allocation efficiency reducing which operations occur (power metric inference, 5.18x), and depth efficiency reducing how many layers each operation uses (early exit, 1.56x). These layers are independent and compound multiplicatively. Applied on top of Saad-Falcon et al.'s 2025 hardware baseline, the combined IPW improvement is estimated at up to approximately 122x versus the 2023 baseline. The full stack including speculative decoding and quality-preserving layers reaches 70x algorithmic improvement alone. Critically, these algorithmic gains are available today on existing hardware — they do not require waiting for the next hardware generation.

Keywords

Artificial intelligence, Machine learning

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average