Large language models store and utilize factual knowledge beyond simple text generation. Understanding where such knowledge is stored within a model is important not only for interpretability, but also for knowledge editing and security analysis. Know...
Large language models store and utilize factual knowledge beyond simple text generation. Understanding where such knowledge is stored within a model is important not only for interpretability, but also for knowledge editing and security analysis. Knowledge editing is based on selectively identifying and modifying specific factual knowledge stored inside a model. Therefore, identifying the layers and modules in which knowledge is concentrated is essential for locating effective editing points and mitigating side effects that may arise during the editing process. In addition, such knowledge storage locations are also important from a security perspective, as they may be exploited for attacks such as malicious factual knowledge injection or backdoor insertion. However, prior studies have primarily focused on GPT-family architectures, and it remains insufficiently understood whether these knowledge storage mechanisms appear similarly across different autoregressive Transformer models. This study quantitatively analyzes the knowledge storage mechanisms of GPT-2-XL, LLaMA-3.2-1B, Qwen-2.5-1.5B, and DeepSeek-R1-Distill-Qwen-1.5B. To this end, we apply three causal analysis methods, Restoration Effect, Severing Effect, and Module Knockout, and compare the structural characteristics through which factual knowledge is stored across architectures. The results show that the storage mechanisms of factual knowledge differ across model architectures. GPT and LLaMA family models tend to store factual knowledge primarily in early MLP layers, whereas Qwen and DeepSeek family models rely more heavily on early Attention layers. In addition, Gini coefficient analysis reveals that these Attention-centered contributions are concentrated in specific layers. Furthermore, Module Knockout provides a more reliable validation of functional importance than Severing Effect alone, particularly in cases where such importance is difficult to confirm through Severing Effect. These findings extend our understanding of interpretability and the feasibility of knowledge editing in large language models and suggest the importance of considering architectural differences in the analysis of security vulnerabilities, including knowledge-editing-based attacks and backdoor insertion points.