科学声明能从大语言模型中移除吗?声明级遗忘的系统性评估
文章背景与核心概要
大语言模型(LM)传统上基于静态的科学语料库进行训练,这使得它们在应对科学知识不断演进(如撤回、被推翻或更新的研究)时面临巨大挑战。尽管机器提供了一种遗忘过时信息同时保留模型效用的方法,但现有方法主要集中在实例级(instance-level)的遗忘,而非相互关联且不断演进的科学声明。
为了弥补这一差距,本文引入了科学声明遗忘(Scientific Claim Unlearning)任务,并推出了一个全新的基准测试——SciUnlearn。作者通过研究表明,当前的遗忘方法未能有效地消除声明级知识,往往只能实现表面的抑制,这凸显了开发专门适用于结构化知识移除技术的必要性。
Summary
大语言模型(LMs)传统上基于静态的科学语料库进行训练,这使得它们在应对科学知识不断演进(如撤回、被推翻或更新的研究)时面临巨大挑战。尽管机器遗忘提供了一种遗忘过时信息同时保留模型效用的方法,但现有方法主要集中在实例级遗忘,而非相互关联、不断演进的科学声明。
Large language models (LMs) are traditionally trained on static scientific corpora, making it challenging to handle the continuous evolution of scientific knowledge (such as retracted, disproven, or updated research). While machine unlearning offers a way to forget obsolete information while preserving model utility, existing methods primarily focus on instance-level forgetting rather than interconnected, evolving scientific claims.
为了弥补这一差距,本文引入了科学声明遗忘任务,并推出了一个全新的基准测试——SciUnlearn。作者证明了当前的遗忘方法无法有效消除声明级知识,通常只能实现表面的抑制,从而凸显了对适合结构化知识移除的专门技术的需求。
To bridge this gap, this paper introduces the task of Scientific Claim Unlearning along with a new benchmark, SciUnlearn. The authors demonstrate that current unlearning methodologies fail to effectively eliminate claim-level knowledge, often achieving merely superficial suppression and underscoring the need for specialized techniques suited for structured knowledge removal.
Metadata & Document Information
- arXiv ID: arXiv:2608.20960 [cs.AI]
- Subject: Computer Science > Artificial Intelligence (
cs.AI) - Conference Venue: EMNLP 2026 Main Conference
- Submission Date: August 21, 2026
- DOI: 10.48550/arXiv.2608.20960
Authors
- Snigdha Paul
- Manasi Patwardhan
- Arman Cohan
Access & Resources
- Full-Text Options:
- View PDF
- HTML (Experimental)
- TeX Source
- External References & Citations:
- NASA ADS
- Google Scholar
- Semantic Scholar
