大语言模型还能解释自己吗?量化对自解释影响的研究
文章背景与核心概要
量化技术已成为加速大语言模型(LLM)推理并降低部署成本的主流手段。然而,量化对模型“自解释”(Self-Explanations, SE)——即模型为自身决策提供推理依据的能力——的影响,此前一直缺乏深入研究。
本文探讨了模型量化如何影响两种主要的自解释形式:自然语言解释(NLE)和反事实示例。研究通过对三种常见量化技术在不同位宽下的表现进行评估,发现量化会导致自解释的质量和忠实度出现轻微下降,但在用户感知层面的连贯性和信任度上,降幅则更为显著。
研究结论指出,虽然量化是实用的压缩策略,但其对解释能力的影响因上下文而异。对于高风险的透明度应用场景,必须对量化后的模型进行实证验证,以确保其解释的可靠性。
📋 摘要
量化是加速推理和减少大语言模型(LLM)部署占用空间的广泛采用的技术。然而,它对自解释(SEs)——即模型为证明自身决策合理性而生成的推理——的影响在很大程度上仍未被探索。
Quantization is a widely adopted technique for accelerating inference and reducing the deployment footprint of Large Language Models (LLMs). However, its influence on Self-Explanations (SEs)—reasoning generated by models to justify their own decisions—has remained largely unexplored.
本文研究了模型量化如何影响两种不同类型的自解释:自然语言解释(NLEs)和反事实示例。通过评估在各种位宽下使用三种常见技术量化的模型,本研究得出了以下关键见解: * 质量与忠实度的适度下降: 量化通常会导致 SE 质量(最高 4.4%)和忠实度(最高 3.9%)的轻微下降。 * 用户信任与连贯性: 用户评估显示,量化可使 SE 的连贯性和可信度显著降低,降幅高达 8.5%。 * 模型规模动态: 在 SE 质量方面,大型模型对量化的抵御能力有限,尽管它们在忠实度上保持了比小型模型更高的水平。 * 没有通用的技术: 没有单一的量化方法能在任务准确性、SE 质量和忠实度方面始终表现出色。
This paper investigates how model quantization affects two distinct types of self-explanations: natural language explanations (NLEs) and counterfactual examples. Evaluating models quantized using three common techniques across various bit widths, the study delivers the following key insights: * Moderate Quality & Faithfulness Drops: Quantization typically causes mild declines in SE quality (up to 4.4%) and faithfulness (up to 3.9%). * User Trust & Coherence: User evaluations show that quantization can considerably diminish both the coherence and trustworthiness of SEs by up to 8.5%. * Model Scale Dynamics: Larger models exhibit limited resilience to quantization regarding SE quality, though they maintain higher levels of faithfulness compared to smaller models. * No Universal Technique: No single quantization method consistently excels across task accuracy, SE quality, and faithfulness.
最终,虽然量化仍然是一种实用的压缩策略,但其影响因上下文而异,这使得对高风险透明度应用进行自解释的实证验证变得至关重要。
Ultimately, while quantization remains a practical compression strategy, its impact varies significantly by context, making empirical validation of self-explanations essential for high-stakes transparency applications.
👥 作者
- Qianli Wang
- Nils Feldhus
- Pepa Atanasova
- Fedor Splitt
- Simon Ostermann
- Sebastian Möller
- Vera Schmitt
- Qianli Wang
- Nils Feldhus
- Pepa Atanasova
- Fedor Splitt
- Simon Ostermann
- Sebastian Möller
- Vera Schmitt
📚 元数据与提交详情
- 学科: 计算与语言 (
cs.CL);人工智能 (cs.AI);机器学习 (cs.LG) - 提交历史:
- v1: 2026年1月1日
- v2 (最新): 2026年8月25日
- 全文与资源:
- 查看 PDF
- HTML 版本 (实验性)
- arXiv DOI
- Subjects: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)- Submission History:
- v1: January 1, 2026
- v2 (Latest): August 25, 2026
- Full-Text & Resources:
- View PDF
- HTML Version (Experimental)
- arXiv DOI
🔗 外部与文献工具
- 引用: Google Scholar | Semantic Scholar | NASA ADS
- 代码与数据发现: Hugging Face | CatalyzeX Code Finder | DagsHub
- Citations: Google Scholar | Semantic Scholar | NASA ADS
- Code & Data Discovery: Hugging Face | CatalyzeX Code Finder | DagsHub
