Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint
Data sources: ZENODO
addClaim

Spanda: Zero-Cost Lexical Entropy Matches Neural Semantic Uncertainty—Until Frontier Models Break It

Authors: Nayak, Bhupen;

Spanda: Zero-Cost Lexical Entropy Matches Neural Semantic Uncertainty—Until Frontier Models Break It

Abstract

Detecting hallucinations in Large Language Models (LLMs) requires estimating uncertainty. Neural Semantic Entropy accurately detects hallucinations by clustering generated paths using an auxiliary Natural Language Inference (NLI) model, but incurs massive computational overhead. In this paper, we propose Spanda, introducing a zero-cost normalized lexical metric ($R_{sc}$). We empirically evaluate models across four scales (1.5B to 120B). We demonstrate that for mid-sized models (7B–27B), $R_{sc}$ achieves an AUROC of 0.889, matching neural semantic entropy without the NLI overhead. However, on frontier models (120B+), we discover that intense RLHF alignment induces 'Confident Mode Collapse'—the model hallucinates the exact same incorrect answer across all paths, causing AUROC to invert to 0.091 and bypassing self-consistency assumptions.

Powered by OpenAIRE graph
Found an issue? Give us feedback