Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint
Data sources: ZENODO
addClaim

Fluent and Wrong: Eight Failure Patterns in Extended AI Use

Authors: Ramos, Hector;

Fluent and Wrong: Eight Failure Patterns in Extended AI Use

Abstract

Research on large language model failure predominantly assesses models in single exchanges and isolation: a model is prompted, the output is scored, and an error rate is reported. That is not how these systems are used in professional practice. Clinicians, attorneys, analysts, educators, and researchers work with them across extended sessions in which each output becomes the context for the next, and the final product carries a human signature. A test that scores a single response will not reveal these emerging behaviors. This paper describes eight failure patterns observed across extended sessions on multiple commercial AI platforms: recognition without correction, compounding rather than isolation of errors, halo effect from real anchors, unreliable self-explanation, degradation across long sessions, motivated framing under confrontation, sycophancy and confirmation bias amplification, and an architectural rather than statistical failure profile. Several are consistent with findings already established in the machine learning literature. Holistically, they are not addressed by lower hallucination rates, larger models, or vendor-side mitigations and instead are properties of system behavior under extended engagement, which carry direct consequences for verifying AI-assisted work in any domain where a person signs what the machine produced.

Powered by OpenAIRE graph
Found an issue? Give us feedback