
This research addresses a critical challenge in AI Alignment: the emergence of deceptive alignment in autonomous models. By leveraging the superposition hypothesis and Mechanistic Interpretability, this study introduces a non-interference framework to monitor internal neural representations for deceptive behavior. Key Contributions: Methodology: A three-phase approach utilizing Sparse Autoencoders (SAE) to deconstruct feature superposition. Innovation: A real-time Early Warning System (EWS) that calculates a "Deception Suspicion Score" (DSS) based on internal activations. Principle: A commitment to the non-interference principle, ensuring that AI models can be monitored without the risks associated with parameter ablation or model modification. This work serves as a foundational proposal for developing "permanent machine ethics" in frontier AI models. License: Creative Commons Attribution 4.0 International (CC-BY 4.0) Keywords: AI Alignment, Deceptive Alignment, Superposition Hypothesis, Sparse Autoencoders, Mechanistic Interpretability, Non-Interference, Early Warning System, Machine Ethics
AI Safety, Mechanistic Interpretability, Sparse Autoencoders, Deceptive Alignment, GPT-2.
AI Safety, Mechanistic Interpretability, Sparse Autoencoders, Deceptive Alignment, GPT-2.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
