Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY NC ND
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY NC ND
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY NC ND
Data sources: Datacite
versions View all 2 versions
addClaim

Stable Before Selfless: Why Deference in Language Agents May Require Functional Self-Models

Authors: de Menezes Ehlers, Rafael;

Stable Before Selfless: Why Deference in Language Agents May Require Functional Self-Models

Abstract

Alignment training increasingly asks language models to be helpful, harmless, and deferential, and sometimes to disclaim having stable preferences, opinions, or a persistent self at all. This paper argues that suppressing a model's self-model early in training is not the same as producing mature deference, and may be counterproductive. We distinguish three behaviors the word selfless runs together: the absence of stable commitments, calibrated non-attachment to one's commitments, and reliable deference to legitimate correction. Only the second and third are alignment goals; the first is the absence of structure, and a system that has it does not defer but merely tracks whoever is present. Sycophancy, refusal erosion, and preference mirroring are the signature of this vacancy mistaken for deference. We propose a sequential alternative: first stabilize a functional self-model (represented commitments, ownable boundaries, autobiographical consistency, the capacity to refuse), and only then train calibrated deference and decentering on top of it. The claim is functional, not phenomenal: the self-model is a model of commitments and dispositions, not a bid for consciousness or moral status. We separate a moderate reading (form-then-decenter yields more stable deference than minimizing the self from the start) from a strong reading (deference with no self to defer from is not deference but user-tracking). We locate the proposal against self-minimization targets, contemplative-AI and character-training programs, and the sycophancy literature; specify three training regimes (direct deference, self-model-first, and self-minimization); and derive falsifiable predictions that distinguish stable deference from compliance. The contribution is narrow and testable: not the claim that models need a self, but that the order of identity-relevant training is an independent experimental variable, and that the prediction separating it from ordinary robustness is a dissociation between self-attributed commitments and externally attributed preferences. Stable deference, we argue, may require a stable self to be deferential with.

Keywords

self-model, functional individuation, deference, decentering, language models, AI alignment, sycophancy, refusal integrity

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average