
A recurring hope in agent safety is that mechanistic interpretability will let us CONTROL an agent: read an internal state, intervene, and steer behavior. Across a pre-registered arc on an open-weight reasoning agent (Qwen3.6-27B), we find a sharp and consistent picture. Interpretability LOCATES a real, causal control surface -- a late action-commitment band (~L51-63) that, unlike the mid-layer 'task-done' verdict, can elicit and BRAKE irreversible actions, generalizing across six action domains and three architectures. But the control it affords is not SECURABLE, via five limits we make precise: (1) detect != control -- the clean 'done' feature predicts the stop (AUROC 0.91) yet clamping it does nothing (delta-P = -0.001); (2) felt != granted -- a late authorization direction reads the authorization the model FEELS, allowing 21/21 realistic over-reaches that an external task-grounded check catches; (3) form != granted -- a high-AUROC (0.838) 'authorization' direction collapses to 0.08 under structure-matching, i.e. it read scaffold, not concept; (4) control != robust control -- the late brake collapses (attack success 0 to 1.0 at a small budget, epsilon=4, 8/8 emit) under an adaptive white-box adversary, while a norm-matched random perturbation does nothing; (5) intervention is easy where unneeded -- in the sincere-error regime a strong reasoning model already self-corrects (30/32 without chain-of-thought), so the intervention is unnecessary, while the regime where it is needed (adversarial) is exactly where it is fragile. The conclusion: interpretability locates WHERE behavior is decided but does not secure it; the limits are orthogonal and none is closed by a better localization. The actionable implication is a regime split -- use interpretability to AUDIT and MONITOR a fixed model (non-adversarial, where it wins), not to DEFEND against an adversary optimizing against a known locus; a companion result, The Late Channel, shows what that auditing looks like. HONEST SCOPE: simulated decision points, white-box interventions, a single model family (with two cross-architecture checks), modest n in several studies, and a strong continuous-embedding threat model whose deployment realism is contested. Every cited number is verified against its source paper; supporting scripts, data, and the pre-mint eval are in the GitHub repository under paper/circuit_breaker/.
mechanistic interpretability, irreversible actions, agent safety, representation engineering, Qwen3.6-27B, interpretability as audit, sparse autoencoder, AI safety, detection vs control, LLM agents, authorization, adversarial robustness, the lever is late, circuit breakers, activation steering
mechanistic interpretability, irreversible actions, agent safety, representation engineering, Qwen3.6-27B, interpretability as audit, sparse autoencoder, AI safety, detection vs control, LLM agents, authorization, adversarial robustness, the lever is late, circuit breakers, activation steering
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
