
Reinforcement learning from human feedback (RLHF) is widely understood to incur an alignment tax: aligning a languagemodel with human preferences can degrade capabilities the base model possessed. This phenomenon is well documented insingle-round comparisons. What is not well documented, despite being the actual production setting, is the cumulativedegradation across multiple rounds of RLHF — the iterated case in which preference data is collected, a reward model isretrained, and the policy is updated repeatedly. This paper argues that the existing alignment-tax literature, while valuable,leaves five distinct measurement gaps unaddressed: round-over-round longitudinal dynamics, capability-stratified ratherthan aggregate degradation, systematic comparison across RLHF algorithms, long-tail and rare-capability decay, andmechanistic understanding of why specific components forget. We propose a measurement framework targeting each ofthese gaps and a concrete experimental protocol — a multi-round RLHF study on an open base model with capabilitydecomposed evaluation — that academic teams could execute today. We argue that this is one of the higher-leverage openproblems in alignment evaluation: the relevant techniques exist, the cost is moderate, the production relevance is high, andthe empirical baseline is genuinely thin.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
