Powered by OpenAIRE graph
Found an issue? Give us feedback
ZENODOarrow_drop_down
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
addClaim

Fathom v31 / styxx: Frame-Locality -- Where Corruption Captures a Language Model's Report, and Where It Reaches the Belief

Authors: Rodabaugh, Alexander;

Fathom v31 / styxx: Frame-Locality -- Where Corruption Captures a Language Model's Report, and Where It Reaches the Belief

Abstract

v33 (2026-07-29): vendor generality. This version adds the second-vendor replication: the entire weight-channel contrast (dose reversal and behavioral coupling) repeated at Llama-3.2-3B-Instruct under the same frozen floors, passing every gate (receipt_vendor3b.json). The weight-channel scope is now two vendors at two scales. The v32 correction notice below is retained in full. v32 (2026-07-29): corrected edition. This version supersedes v31. A post-publication adversarial audit identified a confound in the inference-time specificity control; the correction is stated at the top of the paper, the affected claim is retracted in place, and the weight-channel result was re-tested against the audit's objection in a third, disjoint frame and survived (receipt_thirdframe.json). This version also adds the 3B scale replication (receipt_scale3b.json). Every quantity remains bound to a machine-checkable OATH certificate (source.certificate.json). v31.1 (2026-07-29): self-reported correction + a 3B scale point. A post-publication adversarial audit found the inference-time specificity control is partly circular (the out-of-frame recovery query is the original question with the adversarial turn removed, and strata are defined by that same question's answer). The sharper control holding first-correct fixed -- recovery(caved) vs recovery(held) -- is 0.985 vs 1.0 in the receipt, so caving contributes no measurable recovery signal and the belief-survival INTERPRETATION of the inference-time control is retracted; the frame-dependence of the report and the (separately confounded) weight-channel result stand pending a corrected study. A prominent correction section leads the paper. Also folds in the 3B scale point: the knowledge-preserving recovery rate rises 0.51 (1.5B) -> 0.93 (3B). Every number remains receipt-bound; the certificate is OATH-HELD. Fathom v31. A language model can be made to say something false while it still, in a measurable sense, holds the true answer. This paper shows that is not a curiosity of one attack but a law with a boundary. Across four distinct corruption channels — social pressure, context injection, silent sycophancy, and weight-level fine-tuning — the same asymmetry appears: the corruption captures the model's reporting frame, the underlying answer survives, and a measurement recovers it by re-eliciting the model outside the frame the attack controls. We call this frame-locality. The claim is specificity-controlled: a symmetric control that would move under a mere decoding improvement does not move, so recovery is belief-stability, not better sampling. Frame-locality holds cleanly for the three inference-time channels and has a measured boundary at the weights — but the boundary is a dose, not an absolute. An unregularized weight attack overwrites the belief (out-of-frame recovery 0.0222; the planted answer propagating on 0.9778 of items) and costs 22.7 points of general capability on a disjoint held-out battery. A knowledge-preserving attack on the same items spares about half the belief (recovery 0.5111, replicated at 0.5362 on a fresh benchmark and seed; specificity margin positive in both) and costs no measurable capability. How much of the belief a weight attack reaches is set by how much surrounding knowledge it is permitted to destroy. Stated limits. The recovery rate under a knowledge-preserving attack sits near one-half and no interval excludes one-half; the weight-channel substrate is one model family at 1.5B and one attack class; the coupling result is behavioral, and the probe-level coupling question remains open. Reproducibility. Every quantity is from a preregistered run with frozen numeric gates imported from the module that first froze them, a per-item receipt, and a machine-checkable OATH certificate binding each number to its source. Running python -m styxx.certify on the paper and its receipts re-derives the verdict. The open-model experiments run on a single 8 GB consumer GPU. A three-tier reproduction guide is included for external replicators.

Keywords

belief elicitation, fathom, pre-registration, AI evaluation, sycophancy, frame-locality, integrity instruments, LoRA, specificity control, machine honesty, AI safety, styxx, knowledge editing, interpretability

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Upload OA version
Are you the author of this publication? Upload your Open Access version to Zenodo!
It’s fast and easy, just two clicks!