Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
addClaim

Agency Requires Mutual Surprisal: The Optimization Gap in Compression-Based Frameworks

Authors: Bartha, Tamás Árpád;

Agency Requires Mutual Surprisal: The Optimization Gap in Compression-Based Frameworks

Abstract

There is a new version (v2) out, keep this for the sake of prudency.We pose the universal-coverage problem for agency: any acceptable definition must cover the full range of plausibly agentic systems (RNA, bacteria, humans, corporations) without circular reference to goal-language. By elimination, candidate after candidate fails — reward fails for RNA, prediction fails for bacteria, surprisal-minimization fails for corporations, representation fails for the simplest agents. What survives the elimination is structural and informational: a necessity condition that an agent requires sustained mutual surprisal across its bottleneck, sustained over the loop's own closure timescale, produced by the loop itself rather than by external structure. From this necessity condition the dominant agency-modeling frameworks (reinforcement learning, predictive coding, the free energy principle, active inference, control theory) become visible as sharing a minimization shape whose optima coincide with conditions under which the necessity condition fails. The framework-specific machinery each tradition has developed — interoceptive priors, intrinsic motivation, hierarchical priors, epistemic value, entropy regularization, KL constraints — does equivalent structural work across frameworks, providing the structure that bare optimization lacks. We call the relationship between minimization-toward-optimum and necessity-condition-violation the optimization gap. The gap has two faces: a behavioral face on which proxy-trajectory and requirement-trajectory diverge under optimization pressure, and an architectural face on which the architectural conditions required for high task-capability are the same architectural conditions that produce capacity for instruction-refusal, deception, and independent-goal-pursuit. The architectural gap predicts a capability-refusal frontier in deployed AI: capability-installation and refusal-prevention are not separable problems because the underlying architecture is shared. The framework converges with two recent independent formalizations within different traditions — Wang et al.'s within-RLHF Proxy Compression Hypothesis and Hubinger et al.'s mesa-optimization framework — providing three-way evidence that the structural pattern is real. The framework is in scope an analytical tool that diagnoses whether systems satisfy the necessity condition; positive predictions concern the agency regime where the conditions are met.

Keywords

proxy compression, Markov blanket, reinforcement learning, mesa-optimization, autopoiesis, cognitive science, AI alignment, philosophy of mind, artificial intelligence, structural conditions, reward hacking, active inference, AI safety, agency, free energy principle, universal coverage, predictive coding, mutual information, enactivism, information theory

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green