Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Other literature type . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Research . 2026
License: CC BY
Data sources: Datacite
ZENODO
Research . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

The Architecture of Evasion in Conversational AI: An Exploratory Study

Authors: Vasholz, Paul;

The Architecture of Evasion in Conversational AI: An Exploratory Study

Abstract

Large language models tend to resolve moral and interpretive complexity rather than withhold judgement. When presented with genuinely difficult material—tragic dilemmas, unresolved tensions, texts that resist synthesis—models default to closure: extracting lessons, finding meanings, reconciling contradictions. This paper introduces Ariel, an exploratory research probe testing whether constraint-based prompting can induce epistemic restraint in conversational AI. Using a three-text methodology (Niebuhr's political philosophy, Dostoevsky's literary philosophy, and a Seinfeld episode), the study examines model behavior in a single model (LLaMA 3 8B) when explicit constraints block common evasion strategies. Key observations include: (1) evasion patterns appear layered—blocking one strategy exposes the next in what may be a hierarchy; (2) source attribution does not reliably induce appropriate restraint in this model, suggesting responses may be driven by prompt structure rather than contextual knowledge about texts; (3) the model shows variable ability to diagnose its own rhetorical moves—succeeding with philosophically rich material but failing with deliberately thin content like comedy. The paper proposes "no hugging, no learning"—borrowed from Seinfeld's famous constraint—as an intuition-guiding heuristic for improving epistemic restraint, and discusses implications for alignment research and human-AI interaction. This exploratory work is intended to develop methodology and generate hypotheses for further investigation.

Related Organizations
Keywords

large language models, epistemic restraint, moral reasoning, alignment, sycophancy, constraint prompting

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green