Powered by OpenAIRE graph
Found an issue? Give us feedback
ZENODOarrow_drop_down
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

Predicting How Transformers Attend, Part III: From Attention to Residual Computation — Separating Transport, Writing, and Commitment

Authors: Carles Marín Muñoz;

Predicting How Transformers Attend, Part III: From Attention to Residual Computation — Separating Transport, Writing, and Commitment

Abstract

Third paper in the “Predicting How Transformers Attend” series (Marín 2026), following Part I (Zenodo doi:10.5281/zenodo.19826343), which introduced the Thermodynamic Attention Framework and the closed-form predictor for the attention-decay exponent γ, and Part II (doi:10.5281/zenodo.19960573), the six-axis decomposition. Parts I and II characterised the distance axis of attention; this paper opens the orthogonal depth axis. Its starting observation: the single interpretability question “is the answer here yet?” silently conflates four distinct statements about the residual stream — that information is present, linearly decodable, written, and used. The paper shows these are separate, independently measurable observables, and decomposes residual computation into three: transport (the mean input–output Jacobian; the Jacobian lens), net injection (direct logit attribution, DLA), and commitment (the depth at which the running decision locks). The DLA decomposition is exact: with the final normalisation frozen the output logit equals the sum of per-layer contributions (validated to 7.6×10−6, norm-aware for LayerNorm and RMSNorm). Empirically the paper reports the only ladder-complete Jacobian-lens readability suite across a Pythia scale ladder (70M–2.8B) with cross-family anchors (Llama-2/3, Qwen2.5/3, GPT-2). The readability knee is a family property (Pythia 0.54–0.75 of depth, Llama 0.81–0.84) and the Jacobian-lens advantage band is strikingly heterogeneous (absent at d_head = 256; up to 68× at layer 1 of Qwen2.5-3B). The paper's pivot is a structural theorem: the advantage band is not a “workspace” of stored content but a property of the transport operator itself (Jℓ = ∏k≥ℓ(I + w′k)), which forces a strong anticorrelation (Spearman −0.80 for net injection, pythia-410m; replicated on pythia-2.8b and Qwen3-1.7b) between Jacobian-lens advantage and net injection. A generic denoised mean operator saturates as an estimator at N ≈ 20 prompts and already beats the per-prompt oracle, so the lens advantage measures generic linear transport, not prompt-specific content. This sharpens — and does not contradict — the causal “global workspace” result of Lindsey et al. (Anthropic, 2026): a lens advantage is a hypothesis generator, not causal evidence. Across families the observables generalise but their chronology does not: a single number, Δprep = τwrite − τdecode, classifies families cleanly (Type A ≥ 0.5 vs Type B ≤ 0.3, no overlap). Honest revisions, each by a pre-registered discriminant design and reported in the same register as surviving results: three hyperparameter predictors of the knee are refuted (the L_crit formula — discriminant pythia-1b, predicted 0.94, observed 0.688; a universal constant ≈0.69, killed by the full ladder; the RoPE fraction, inconclusive by a frozen control rule). L_crit is not discarded — it governs a distinct, causally-validated attention-regime transition, decoupled from the knee. τcommit is architectural, not a deep law. The relation of the chronology to the Part I exponent γ is left inconclusive (two γ-measurement protocols disagree and the direction of any co-ordination flips between them). All experiments are pre-registered with frozen decision rules; the load-bearing algebraic identities are machine-verified in Lean 4/Mathlib in a companion formalization paper (a separate release). English (14 pages) and Spanish (15 pages) editions are included. Prepared with the assistance of an AI system (Claude, Anthropic); the research and all claims are the author's responsibility.

This is Part III of the "Predicting How Transformers Attend" series and the natural evolution of the first two parts — readers new to the series should start there: Part I — Analytic Power-Law Theory (10.5281/zenodo.20314038): explains from first principles why attention weights decay as a power law of distance in RoPE transformers, introducing the Thermodynamic Attention Framework (TAF) and the closed-form predictor γ_Padé(θ,T) for the decay exponent γ. Part II — A Six-Axis Decomposition (10.5281/zenodo.19960573): extends TAF into a six-axis decomposition of γ (including the learned-imprint axis ν≈−1/(2π)), adds a 4-bit NF4 precision-direction rule and a bimodal Hagedorn phase structure, and machine-verifies the algebraic backbone in Sage + Lean. Part III is the natural next step: it moves from the attention map to the residual stream, separating transport, writing, and commitment.English and Spanish editions included. Artifacts — the ladder-complete Jacobian-lens suite (70M–2.8B), all readability curves, the DLA traces, the pre-registrations with dated decision rules, and the metric specification — are in the research repository (part3/experiments/workspace_h1/). Companion formalization paper 'A Machine-Checked Backbone for the TAF' (Lean 4 / Mathlib verification of the load-bearing identities) is a separate later release. Prepared with AI assistance (Claude, Anthropic); the mathematics and all claims are the author's responsibility.

Keywords

LLM, falsification, Pythia, pre-registration, rotary position embedding, direct logit attribution, scaling ladder, Thermodynamic Attention Framework, logit lens, depth axis, attention decay, readability knee, global workspace, transport operator, residual stream, machine verification, mechanistic interpretability, RoPE, deep learning, Jacobian lens, honest revision, transformer, critical exponents, attention mechanism, commitment depth, Lean Mathlib4, residual computation

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Upload OA version
Are you the author of this publication? Upload your Open Access version to Zenodo!
It’s fast and easy, just two clicks!