
We ask whether an LLM agent's action commitment—which tool it calls—routes through the emergent verbalizable ‘global workspace’ (Anthropic, 2026) at the same network depth as its verbalizable answer. Using a row-restricted J-lens estimator whose readout directions are validated by a specificity control, we find a depth lag: the verbalizable direction becomes causally steerable for the tool commitment strictly deeper than for the answer. There is a depth band where steering along the verbalizable direction specifically reroutes a held multi-hop answer but not the committed tool (at or below a magnitude-matched random-direction control); the action becomes steerable via the same direction only deeper. This replicates across two architectures (a dense 27B and an MoE 20B); the absolute onset depths are model-dependent but the answer→action lag is consistent. Ablating the top verbalizable directions at the decision point leaves the commitment intact. A verbalizable (J-lens-style) monitor thus reads a reasoning agent's answer a depth-band before its action is committed. Scoped to two open-weights models; the steer and ablation numbers are independently GPU-reproduced (all six targeted dissociation counts, exactly), and every positive dissociation is significant against its random control (Fisher exact, p from 1.5e-4 to 2.5e-18). Part of the OpenInterpretability arc on long-horizon agent control (beat 11).
J-lens, mechanistic interpretability, tool use, knowledge-action gap, AI safety, LLM agents, global workspace, circuit analysis, activation steering, mixture-of-experts
J-lens, mechanistic interpretability, tool use, knowledge-action gap, AI safety, LLM agents, global workspace, circuit analysis, activation steering, mixture-of-experts
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
