Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

The Verbosity Premium: What RLHF-Induced Token Inflation Costs the AI Industry

Authors: Kearney, John;

The Verbosity Premium: What RLHF-Induced Token Inflation Costs the AI Industry

Abstract

We aggregate published measurements of RLHF-induced response length inflation across the literature and compute the first industry-scale estimate of its economic cost. Alignment training systematically inflates output length: sentences triple after SFT, DPO doubles response length within the first 10% of training, and on one benchmark 98% of PPO reward improvement is attributable to length alone. Verbosity compensation rates range from 13.6% to 74.2% across 14 models, and output tokens cost 4-8x more than input tokens across all frontier providers. Combining published verbosity rates, real-world token volumes, and current API pricing, we estimate the annual verbosity premium at 500M to 1.8B, with a central estimate of 1.2B (approximately 14% of total industry inference spend). We survey 12 training-side mitigations and show that all target response length rather than information density. A 500-token response with 50 atomic facts is efficient; the same length with 10 facts restated five ways is waste. Length penalties cannot distinguish these cases. Drawing on rate-distortion theory and evidence that factual precision degrades with response length, we argue the correct optimization target is information density (supported facts per token) and present two concrete density-aware reward formulations.

Keywords

token efficiency, preference optimization, language model alignment, RLHF, verbosity, information density, DPO, inference cost

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green