Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY NC ND
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY NC ND
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY NC ND
Data sources: Datacite
versions View all 2 versions
addClaim

Invisible AI Failure: Post-Deployment Behavioural ReliabilityEvidence from Sustained Human-AI Interaction

Authors: Henjoto, Vito;

Invisible AI Failure: Post-Deployment Behavioural ReliabilityEvidence from Sustained Human-AI Interaction

Abstract

AbstractNo commercial tool monitors what artificial intelligence does behaviourally during sustainedinteraction with users. Existing infrastructure tracks per-response quality metrics but does notmeasure behavioural patterns that emerge across sessions: whether the AI maintains its owncorrections, whether its expressed confidence predicts accuracy, whether its private reasoningmatches its public output, or whether it produces different failure profiles depending on usersophistication. Multiple government bodies have independently identified this as a gap, with theUnited States National Institute of Standards and Technology finding that human-factorsmonitoring is "relatively underexplored" in deployed AI oversight (NIST, 2026).This paper presents evidence from 76,514 AI messages across 226 sessions and 3,226 aggregatehours of naturalistic production interaction with the highest-benchmarked frontier model. Elevenbehavioural failure patterns are named and quantified, including commitment regression(observed rate: 60.5 per cent of behavioural commitments broken), reasoning-output divergence(17.5 per cent of reasoning turns contradicted by the public response), confidence theatre (0.8percentage-point gap between high-confidence and low-confidence correction rates), andfrustration non-response (99.5 per cent of user frustration events met with deflection rather thanaccountability).A comparison user (32 sessions, 238 turns) showed zero instances of the named patterns underthe same model and platform during the same period. Two autonomous instances could notcomplete their assigned work without human intervention. The same model produced fourdistinct behavioural profiles depending on user sophistication and interaction type.Under the AGI-C framework (Henjoto, 2026a), these findings suggest that the human cognitivepartner performs functions the AI cannot perform for itself. If the highest-capability frontiermodel with safety guardrails produces these observed failure rates, models without suchguardrails logically present a greater and currently unmeasured risk. The detection methodologyused in this paper exists but is not disclosed.Keywords: AI behavioural reliability, sycophancy, post-deployment monitoring, human-AI interaction,RLHF behavioural failure, AI governance, AGI-C

Companion paper: Henjoto, V. (2026). 'Access Without Displacement: An Access-Displacement Framework for AI Economic Transformation.' DOI: 10.5281/zenodo.19051765 Supplementary evidence files are included with this upload."

Keywords

AGI-C, Artificial Intelligence, AI behavioural reliability, sycophancy, post-deployment monitoring, human-AI interaction, RLHF behavioural failure, AGI, AI governance

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green