
arXiv: 2407.01178
The training and inference of large language models (LLMs) are together a costly process that transports knowledge from raw data to meaningful computation. Inspired by the memory hierarchy of the human brain, we reduce this cost by equipping LLMs with explicit memory, a memory format cheaper than model parameters and text retrieval-augmented generation (RAG). Conceptually, with most of its knowledge externalized to explicit memories, the LLM can enjoy a smaller parameter size, training cost, and inference cost, all proportional to the amount of remaining “abstract knowledge”. As a preliminary proof of concept, we train from scratch a 2.4 B LLM, which achieves better performance than much larger LLMs as well as RAG models, and maintains higher decoding speed than RAG. The model is named ${\rm Memory}^3$, since explicit memory is the third form of memory in LLMs after implicit memory (model parameters) and working memory (context key-values). We introduce a memory circuitry theory to support the externalization of knowledge, and present novel techniques including a memory sparsification mechanism that makes storage tractable and a two-stage pretraining scheme that facilitates memory formation.
FOS: Computer and information sciences, Artificial intelligence, Computer Science - Machine Learning, Computer Science - Computation and Language, Computer Science - Artificial Intelligence, AI database, I.2.7, 68T50, large language model, Machine Learning (cs.LG), explicit memory, efficient inference, Artificial Intelligence (cs.AI), large-scale pretraining, Computation and Language (cs.CL)
FOS: Computer and information sciences, Artificial intelligence, Computer Science - Machine Learning, Computer Science - Computation and Language, Computer Science - Artificial Intelligence, AI database, I.2.7, 68T50, large language model, Machine Learning (cs.LG), explicit memory, efficient inference, Artificial Intelligence (cs.AI), large-scale pretraining, Computation and Language (cs.CL)
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 7 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
