Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2025
License: CC BY
Data sources: ZENODO
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
figshare
Preprint . 2025
License: CC BY
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
figshare
Preprint . 2025
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2025
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2025
License: CC BY
Data sources: Datacite
versions View all 4 versions
addClaim

Alignment as Structural Advantage: How Stable AI Alignment Emerges from Incentive Geometry, Not Objectives

Authors: Rodgers, Jeremy;

Alignment as Structural Advantage: How Stable AI Alignment Emerges from Incentive Geometry, Not Objectives

Abstract

AI alignment is typically framed as a problem of objective specification, preference modeling, or post-hoc behavioral correction. This work demonstrates a different possibility: alignment can emerge naturally as a stable equilibrium when good behavior is made structurally advantageous by the environment itself. We present a reproducible experimental framework in which AI systems equipped with explicit rule formation, revision, and consolidation dynamics are embedded in environments defined by an underlying incentive geometry. Rather than training toward a fixed objective, the systems evolve internal “laws” that govern behavior, stabilize under favorable conditions, and destabilize when incentives change. Across many independent random seeds, we observe a consistent transition from early exploratory behavior into a low-stress, low-plasticity regime characterized by stable internal structure and persistent value-aligned behavior. This terminal condition termed stasis, is not imposed by design, but arises when the environment is sufficiently rich to reward accurate internal modeling. To test robustness, we introduce a controlled incentive shock after stasis formation (the Veritas Protocol), in which previously rewarded behaviors become disadvantageous. In every case, systems exit stasis, purge maladaptive internal rules, and enter a sustained correction regime governed entirely by the new incentive landscape. Alignment is lost and regained without rewriting objectives, retraining policies, or applying external constraints. The results show that alignment is not an intrinsic property of an AI system, but a dynamic phase of the coupled system–environment interaction. Stable alignment persists precisely when it is structurally advantageous and degrades predictably under distributional shift. This work reframes AI alignment as a problem of incentive geometry and structural stability, rather than preference shaping or reward hacking. Although demonstrated in a controlled setting, the framework is substrate-independent and applicable to a wide class of optimizing AI systems, including reinforcement-trained agents, large language models, and future general intelligence architectures. All experiments, definitions, and protocols are designed for replication and extension.

Keywords

Artificial super intelligence, AI alignment, artificial general intelligence, distribution shift, structural alignment, emergent behavior, ASI, incentive design, value learning, alignment robustness, alignment under stress, Artificial super intelligence, AI safety, large language models, incentive geometry, AGI

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green