
This draft proposes a geometric view of vanilla transformer layers as discrete GL(d) connections on the token graph when the connection is defined by the cross-Jacobian at frozen activations. Under per-token basis changes Ri ∈ GL(d), the edge maps Γij = ∂Yi/∂Xj transform covariantly (Ri-1ΓijRj). Residual addition acts as a forward-Euler integrator; LayerNorm fixes an affine gauge slice ℝd/(ℝ+ × ℝ·1) ≅ Sd-2. We define discrete curvature and Wilson loops for short cycles, give a multi-head corollary, and contrast the exact Jacobian connection with the common "values-only" proxy. The appendix provides efficient JVP/VJP recipes to probe Γij without forming full Jacobians, plus an experiments template (curvature heatmaps, canonicalized pruning, transport-aware distillation). Key FindingsWe define discrete curvature and Wilson loops to measure the "non-integrability" of the transformer's path. Empirical validation (included) confirms that while gradient connections are flat (||H - I|| < 10-9), Transformer connections exhibit significant curvature (||H - I|| ≈ 8.0), identifying a geometric source of hallucination and path-dependence. AvailabilityThis is a technical report/preprint. Full repository access (JVP/VJP implementation) is available on request. Keywords: transformers; gauge theory; GL(d); Jacobian; LayerNorm; Neural ODE; discrete connection; Wilson loops; token graph; attention mechanisms; mechanistic interpretability; geometric deep learning
mechanistic interpretability, gauge theory, geometric deep learning, GL(d), token graph, transformers, LayerNorm, Wilson loops, Neural ODE, attention mechanism, discrete connection, Jacobian
mechanistic interpretability, gauge theory, geometric deep learning, GL(d), token graph, transformers, LayerNorm, Wilson loops, Neural ODE, attention mechanism, discrete connection, Jacobian
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
