文章背景与核心概要
知识图谱推理(KGR)旨在利用知识图谱(KG)中的结构证据来发现潜在事实。尽管大语言模型(LLM)通过上下文学习在KGR任务中展现出了巨大潜力,但它们常常受到“推理证据感知漂移”问题的困扰——即KG结构上下文与LLM参数化知识之间存在不一致性,这破坏了推理的忠实度和有效性。
为了克服这一挑战,本文提出了结构内化规则语言模型(SIRLM),通过结构化规则生成将结构知识学习与推理逻辑评估紧密结合。该框架引入了结构内化规则生成器(SIRG)、基于结构不变性学习的KG分词器,以及利用规则约束消息传播的神经符号推理机。SIRLM能够无缝集成到标准的LLM训练范式(如SFT和GRPO)中,在36个数据集上对17种最先进的方法展现出了显著的优越性。
Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning
Summary
Structure-Internalized Rule Language Model (SIRLM) is a novel framework designed to improve Knowledge Graph Reasoning (KGR) using Large Language Models (LLMs). While LLMs have shown promise in KGR via in-context learning, they often suffer from reasoning evidence perception drift—an inconsistency between KG structural context and LLM parametric knowledge that undermines reasoning faithfulness and effectiveness.
To overcome this, SIRLM couples structural knowledge learning with reasoning logic evaluation through structural rule generation. The framework introduces: - Structure-Internalized Rule Generator (SIRG): Incorporates an in-context learning block with a structural relation memory to coordinate structural and parametric knowledge. - KG Tokenizer: Built on structural invariance learning to supply learnable structural representations. - Neuro-Symbolic Reasoner: Utilizes rule-constrained message propagation to offer faithful rule-execution feedback.
SIRLM integrates seamlessly into standard LLM training paradigms like SFT and GRPO, demonstrating significant superiority over 17 state-of-the-art methods across 36 datasets.
Paper Metadata
- arXiv Identifier:
arXiv:2608.17443[cs.AI]- Subject: Computer Science > Artificial Intelligence (
cs.AI)- Submission Date: August 18, 2026
- DOI: 10.48550/arXiv.2608.17443
Authors
- Xingrui Zhuo
- Jiapu Wang
- Manzong Huang
- Gongqing Wu
- Xindong Wu
Abstract
知识图谱推理(KGR)旨在利用知识图谱(KG)中可用的结构证据来发现潜在事实,这对KGR模型的结构语义理解能力提出了挑战。近期研究表明,大语言模型(LLM)通过灵活的上下文学习,能够在KGR任务上取得显著进展。然而,KG结构上下文与LLM参数化知识之间固有的表征不一致问题仍未得到充分解决。这一局限性阻碍了LLM有效感知与KG约束相一致的推理证据,从而削弱了推理的有效性和忠实度。我们将此问题称为LLM在知识图谱上的推理证据感知漂移。为了解决这一问题,我们提出了结构内化规则语言模型(SIRLM),它以结构规则生成为核心,将结构知识的参数化学习与推理逻辑的忠实度评估相结合,使LLM能够紧密锚定在以KG为基础的证据上。具体而言,我们首先设计了结构内化规则生成器(SIRG),它包含一个配备了结构关系记忆的上下文学习模块,以协调结构知识和参数化知识。此外,我们为SIRG配备了基于结构不变性学习的KG分词器,以及基于规则约束消息传播的神经符号推理机。这些组件分别赋予了SIRG可学习的结构表征和忠实的规则执行反馈。我们的SIRLM可以无缝集成到标准的LLM训练范式中,例如SFT和GRPO。在36个数据集上与17种最先进的KGR方法进行的广泛实验证明了SIRLM的显著优越性。
Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models. Recent studies have demonstrated that Large Language Models (LLMs) can achieve remarkable progress on KGR tasks via flexible in-context learning. However, the inherent representation inconsistency between KG structural context and LLM parametric knowledge remains inadequately addressed. This limitation prevents LLMs from effectively perceiving reasoning evidence that aligns with KG constraints, which undermines both the effectiveness and faithfulness of reasoning. We refer to this problem as reasoning evidence perception drift of LLMs over KGs. To address this problem, we propose a Structure-Internalized Rule Language Model (SIRLM), which centers on structural rule generation to couple the parametric learning of structural knowledge with the faithfulness evaluation of reasoning logic, enabling LLMs to anchor tightly to KG-grounded evidence. Specifically, we first design a Structure-Internalized Rule Generator (SIRG), which incorporates an in-context learning block augmented with a structural relation memory to coordinate structural and parametric knowledge. Furthermore, we equip SIRG with a KG tokenizer based on structural invariance learning and a neuro-symbolic reasoner based on rule-constrained message propagation. These components provide SIRG with learnable structural representations and faithful rule-execution feedback, respectively. Our SIRLM can be seamlessly integrated into standard LLM training paradigms, such as SFT and GRPO. Extensive experiments against 17 state-of-the-art KGR methods on 36 datasets demonstrate the significant superiority of SIRLM.
Links & Resources
- Full-Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS