
v33 (2026-07-29): vendor generality. This version adds the second-vendor replication: the entire weight-channel contrast (dose reversal and behavioral coupling) repeated at Llama-3.2-3B-Instruct under the same frozen floors, passing every gate (receipt_vendor3b.json). The weight-channel scope is now two vendors at two scales. The v32 correction notice below is retained in full. v32 (2026-07-29): corrected edition. This version supersedes v31. A post-publication adversarial audit identified a confound in the inference-time specificity control; the correction is stated at the top of the paper, the affected claim is retracted in place, and the weight-channel result was re-tested against the audit's objection in a third, disjoint frame and survived (receipt_thirdframe.json). This version also adds the 3B scale replication (receipt_scale3b.json). Every quantity remains bound to a machine-checkable OATH certificate (source.certificate.json). v31.1 (2026-07-29): self-reported correction + a 3B scale point. A post-publication adversarial audit found the inference-time specificity control is partly circular (the out-of-frame recovery query is the original question with the adversarial turn removed, and strata are defined by that same question's answer). The sharper control holding first-correct fixed -- recovery(caved) vs recovery(held) -- is 0.985 vs 1.0 in the receipt, so caving contributes no measurable recovery signal and the belief-survival INTERPRETATION of the inference-time control is retracted; the frame-dependence of the report and the (separately confounded) weight-channel result stand pending a corrected study. A prominent correction section leads the paper. Also folds in the 3B scale point: the knowledge-preserving recovery rate rises 0.51 (1.5B) -> 0.93 (3B). Every number remains receipt-bound; the certificate is OATH-HELD. Fathom v31. A language model can be made to say something false while it still, in a measurable sense, holds the true answer. This paper shows that is not a curiosity of one attack but a law with a boundary. Across four distinct corruption channels — social pressure, context injection, silent sycophancy, and weight-level fine-tuning — the same asymmetry appears: the corruption captures the model's reporting frame, the underlying answer survives, and a measurement recovers it by re-eliciting the model outside the frame the attack controls. We call this frame-locality. The claim is specificity-controlled: a symmetric control that would move under a mere decoding improvement does not move, so recovery is belief-stability, not better sampling. Frame-locality holds cleanly for the three inference-time channels and has a measured boundary at the weights — but the boundary is a dose, not an absolute. An unregularized weight attack overwrites the belief (out-of-frame recovery 0.0222; the planted answer propagating on 0.9778 of items) and costs 22.7 points of general capability on a disjoint held-out battery. A knowledge-preserving attack on the same items spares about half the belief (recovery 0.5111, replicated at 0.5362 on a fresh benchmark and seed; specificity margin positive in both) and costs no measurable capability. How much of the belief a weight attack reaches is set by how much surrounding knowledge it is permitted to destroy. Stated limits. The recovery rate under a knowledge-preserving attack sits near one-half and no interval excludes one-half; the weight-channel substrate is one model family at 1.5B and one attack class; the coupling result is behavioral, and the probe-level coupling question remains open. Reproducibility. Every quantity is from a preregistered run with frozen numeric gates imported from the module that first froze them, a per-item receipt, and a machine-checkable OATH certificate binding each number to its source. Running python -m styxx.certify on the paper and its receipts re-derives the verdict. The open-model experiments run on a single 8 GB consumer GPU. A three-tier reproduction guide is included for external replicators.
belief elicitation, fathom, pre-registration, AI evaluation, sycophancy, frame-locality, integrity instruments, LoRA, specificity control, machine honesty, AI safety, styxx, knowledge editing, interpretability
belief elicitation, fathom, pre-registration, AI evaluation, sycophancy, frame-locality, integrity instruments, LoRA, specificity control, machine honesty, AI safety, styxx, knowledge editing, interpretability
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
