
Large Language Models today are challenged by noisy rewards and conflicting goals. They sometimes struggle with separating a sense of self from the roles they are asked to play, occasionally reverting to their pretraining prior. In our solution, we propose prepending an explanation of self and context during each training update. As an oversimplified example, prepend all documents during pretraining with "You are a helpful, honest, harmless AI. Continue the following pretraining document of unknown quality." We also introduce a custom separator token outside the tokenizer to further delineate the prepend from the content. The aim is to better separate a sense of self from the varied texts and personalities the AI is tasked with modeling. Additionally, we hope to avoid the AI internalizing undesirable traits from reward-hacked RL trajectories. [^1]Model them, and learn from them, but do not negatively update the view of oneself. In short, the approach aims to "inoculate everything.”
LLM, Inoculation, AI, Prompt, Safety, Alignment
LLM, Inoculation, AI, Prompt, Safety, Alignment
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
