从大型语言模型中直接构建消歧知识库
文章背景与核心概要
自动化知识库构建(AKBC)是自然语言处理(NLP)领域的经典任务,传统方法通常依赖外部文本语料库或以维基百科为中心的结构。近年来,研究人员尝试将大型语言模型(LLM)直接作为知识源来生成知识库。然而,由于LLM本身缺乏对实体的原生表征,这些方法往往会产生重复条目和概念混淆的问题。
本文介绍了 GPTKB 2.0,这是一种旨在直接从LLM中构建消歧知识库的新型方法,同时兼顾了可扩展性和高消歧准确率。通过在生成过程中对实体、关系和类别进行实时消歧,作者成功构建了一个包含超过 100万个已消歧实体 和 3840万个三元组 的百万级知识库,标志着首个具有显式内部规范化的LLM原生知识库的诞生。
摘要 (Abstract)
Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy. We analyze the central design decisions and characterize the trade-offs between accuracy, scale, and cost. We execute GPTKB 2.0 at scale, obtaining a materialized KB containing over 1M disambiguated entities and 38.4M triples. This represents the first million-scale LLM-native KB with explicit internal canonicalization of entities, relations, and classes, a significant departure from prior Wikimedia-centric works.
自动化知识库构建(AKBC)是自然语言处理(NLP)的核心任务,近期的研究提出直接从大型语言模型(LLM)生成知识库,将模型本身作为知识源。然而,LLM原生并不具备实体的表征,这会导致重复条目和概念混淆。我们提出了 GPTKB 2.0,这是一种直接从LLM构建消歧知识库的方法。GPTKB 2.0 结合了对实体、关系和类别的实时消歧,经过精心设计以同时满足可扩展性和消歧准确率的要求。我们分析了核心设计决策,并评估了准确率、规模和成本之间的权衡。我们大规模运行了 GPTKB 2.0,获得了一个包含超过100万个消歧实体和3840万个三元组的实体化知识库。这是首个对实体、关系和类别进行显式内部规范化的百万级LLM原生知识库,与先前以维基百科为中心的工作有着显著的不同。
元数据与出版详情 (Metadata & Publication Details)
| 字段 (Field) | 详情 (Details) |
|---|---|
| arXiv 标识符 | arXiv:2608.03729 [cs.CL] |
| DOI | 10.48550/arXiv.2608.03729 |
| 一级学科 | 计算与语言 (cs.CL) |
| 二级学科 | 人工智能 (cs.AI)、数据库 (cs.DB) |
| 作者 | Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski |
| 提交时间线 | 2026年8月4日提交;2026年9月2日最后修订(版本3) |
| 项目网站 | gptkb.org |
Field Details arXiv Identifier arXiv:2608.03729 [cs.CL] DOI 10.48550/arXiv.2608.03729 Primary Subject Computation and Language ( cs.CL)Secondary Subjects Artificial Intelligence ( cs.AI), Databases (cs.DB)Authors Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski Submission Timeline Submitted on 4 Aug 2026; Last revised 2 Sep 2026 (Version 3) Project Website gptkb.org
资源与访问链接 (Resources & Access Links)
- 全文选项:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 许可证: 知识共享署名 4.0 国际许可协议
- 外部引用与工具:
- Google 学术
- Semantic Scholar
- NASA ADS
- Full-Text Options:
- View PDF
- HTML Version (Experimental)
- TeX Source
- License: Creative Commons Attribution 4.0 International
- External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS