Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
addClaim

Wisdom Is Not Capability: A Self-Validating Measurement Standard for Grounding-Limited AI Agents

Authors: Zhang, Mian;

Wisdom Is Not Capability: A Self-Validating Measurement Standard for Grounding-Limited AI Agents

Abstract

The dominant regime of AI evaluation mistakes capability for wisdom. Benchmarks measure an intercept: how well a system performs when the task is framed, evidence is fresh, and the world has not yet moved. Deployment asks for a slope: whether the system remains grounded as the world drifts, evidence decays, feedback becomes sparse, and reasoning elaborates beyond what reality supports. This paper defines wisdom as maintenance of grounding under drift and introduces effective grounding, phi_eff, as an operational quantity for measuring how much judgment remains closed to external evidence after capability-driven self-elaboration. The central prediction is the capability trap: when grounding is the bottleneck, adding reasoning or compute cannot create signal; it either leaves discrimination unchanged or makes the system more confidently wrong. The paper builds a four-part black-box instrument suite: a wisdom instrument, a grounding-gate audit, a controlled marginal-edge evaluator, and an N4 reasoning-ablation judge. Each instrument is self-validated on known objects before judging a real system. The suite is applied to Soul OS, a real shadow-only proof-carrying trading agent. Across a 24,000-event grounding-gate audit, a 1,316-row controlled edge test, a 414,813-row residual ledger, a 40-row exact live-smoke measurement, and a 40-pair DeepSeek V4 reasoning ablation, the system produces a clean null: no controlled edge, no demonstrable per-decision discrimination, and no rescue from higher reasoning. N4 has 40/40 complete LOW/HIGH pairs with zero pair-integrity failures; high reasoning lowers AUC and worsens Brier in point estimate, but the sample is not powered for a fully significant harm claim. The defensible result is sharper: neither reasoning level shows demonstrable discrimination; higher reasoning did not help and trended toward harm. This is not an alpha claim, trading claim, or live-deployment claim. It is a measurement-standard paper: a ruler strict enough to say when reasoning helps, when reasoning hurts, when evidence is insufficient, and when the correct action is not to act.

Keywords

measurement standard, capability, AI evaluation, deployed agents, reasoning ablation, negative results, grounding, calibration, proof-carrying AI, evidence gate, capability trap, wisdom

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average