跳转至

大语言模型还能解释自己吗?量化对自解释影响的研究

文章背景与核心概要

量化技术已成为加速大语言模型(LLM)推理并降低部署成本的主流手段。然而,量化对模型“自解释”(Self-Explanations, SE)——即模型为自身决策提供推理依据的能力——的影响,此前一直缺乏深入研究。

本文探讨了模型量化如何影响两种主要的自解释形式:自然语言解释(NLE)和反事实示例。研究通过对三种常见量化技术在不同位宽下的表现进行评估,发现量化会导致自解释的质量和忠实度出现轻微下降,但在用户感知层面的连贯性和信任度上,降幅则更为显著。

研究结论指出,虽然量化是实用的压缩策略,但其对解释能力的影响因上下文而异。对于高风险的透明度应用场景,必须对量化后的模型进行实证验证,以确保其解释的可靠性。


📋 摘要

量化是加速推理和减少大语言模型(LLM)部署占用空间的广泛采用的技术。然而,它对自解释(SEs)——即模型为证明自身决策合理性而生成的推理——的影响在很大程度上仍未被探索。

Quantization is a widely adopted technique for accelerating inference and reducing the deployment footprint of Large Language Models (LLMs). However, its influence on Self-Explanations (SEs)—reasoning generated by models to justify their own decisions—has remained largely unexplored.

本文研究了模型量化如何影响两种不同类型的自解释:自然语言解释(NLEs)反事实示例。通过评估在各种位宽下使用三种常见技术量化的模型,本研究得出了以下关键见解: * 质量与忠实度的适度下降: 量化通常会导致 SE 质量(最高 4.4%)和忠实度(最高 3.9%)的轻微下降。 * 用户信任与连贯性: 用户评估显示,量化可使 SE 的连贯性和可信度显著降低,降幅高达 8.5%。 * 模型规模动态: 在 SE 质量方面,大型模型对量化的抵御能力有限,尽管它们在忠实度上保持了比小型模型更高的水平。 * 没有通用的技术: 没有单一的量化方法能在任务准确性、SE 质量和忠实度方面始终表现出色。

This paper investigates how model quantization affects two distinct types of self-explanations: natural language explanations (NLEs) and counterfactual examples. Evaluating models quantized using three common techniques across various bit widths, the study delivers the following key insights: * Moderate Quality & Faithfulness Drops: Quantization typically causes mild declines in SE quality (up to 4.4%) and faithfulness (up to 3.9%). * User Trust & Coherence: User evaluations show that quantization can considerably diminish both the coherence and trustworthiness of SEs by up to 8.5%. * Model Scale Dynamics: Larger models exhibit limited resilience to quantization regarding SE quality, though they maintain higher levels of faithfulness compared to smaller models. * No Universal Technique: No single quantization method consistently excels across task accuracy, SE quality, and faithfulness.

最终,虽然量化仍然是一种实用的压缩策略,但其影响因上下文而异,这使得对高风险透明度应用进行自解释的实证验证变得至关重要。

Ultimately, while quantization remains a practical compression strategy, its impact varies significantly by context, making empirical validation of self-explanations essential for high-stakes transparency applications.


👥 作者

  • Qianli Wang
  • Nils Feldhus
  • Pepa Atanasova
  • Fedor Splitt
  • Simon Ostermann
  • Sebastian Möller
  • Vera Schmitt
  • Qianli Wang
  • Nils Feldhus
  • Pepa Atanasova
  • Fedor Splitt
  • Simon Ostermann
  • Sebastian Möller
  • Vera Schmitt

📚 元数据与提交详情

  • 学科: 计算与语言 (cs.CL);人工智能 (cs.AI);机器学习 (cs.LG)
  • 提交历史:
  • v1: 2026年1月1日
  • v2 (最新): 2026年8月25日
  • 全文与资源:
  • 查看 PDF
  • HTML 版本 (实验性)
  • arXiv DOI
  • Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
  • Submission History:
  • v1: January 1, 2026
  • v2 (Latest): August 25, 2026
  • Full-Text & Resources:
  • View PDF
  • HTML Version (Experimental)
  • arXiv DOI

🔗 外部与文献工具


许可协议:知识共享署名 4.0 国际许可协议
license icon

License: Creative Commons Attribution 4.0 International
license icon