可扩展的克罗内克-费雪近似:针对十亿参数语言模型压缩的高效黑塞矩阵分析
文章背景与核心概要
在大模型压缩与优化领域,对包含数十亿参数的网络进行精确的黑塞矩阵(Hessian)分析长期以来由于计算和存储开销过大而难以实现。传统的全矩阵计算方法在面对超大规模参数时完全不可行,而忽略跨层交互的简化方法又往往无法捕捉真实的模型敏感性。
本文提出了一种可扩展的基于克罗内克积(Kronecker-based)的近似方法,能够在无需存储整个费雪信息矩阵(Fisher matrix)的前提下捕获跨层交互。通过该框架,研究人员首次在多个主流大模型家族中揭示了普遍存在的脆弱性模式,例如值投影层(value projection layers)表现出最高的敏感性和最强的跨层相关性。该研究在量化、稀疏化、层间损坏及微调等大量实验中得到验证,为指导模型压缩、混合精度分配以及自适应低秩分解提供了兼具理论基础与实用价值的强大工具。
📌 摘要 (Summary)
This paper introduces a scalable Kronecker-based approximation method that captures cross-layer interactions without the need to store the entire Fisher matrix. This breakthrough enables practical Hessian analysis for billion-parameter networks, where full matrix computation is typically infeasible. The authors uncover consistent vulnerability patterns across major model families—most notably that value projection layers display the highest sensitivity and strongest cross-layer correlations. Validated through extensive experiments in quantization, sparsification, inter-layer corruption, and fine-tuning, the proposed framework offers a theoretically grounded tool for guided model compression and optimization.
本文引入了一种可扩展的基于克罗内克积的近似方法,该方法能够在无需存储整个费雪矩阵的情况下捕获跨层交互。这一突破使得针对十亿参数网络的实用黑塞矩阵分析成为可能(在此之前,完整矩阵计算通常是不可行的)。作者揭示了各大模型家族中一致的脆弱性模式——最显著的是,值投影层(value projection layers)表现出最高的敏感性和最强的跨层相关性。通过在量化、稀疏化、层间损坏和微调方面的广泛实验验证,该提出的框架为指导模型压缩和优化提供了一个有理论基础的工具。
📄 文章元数据 (Article Metadata)
- arXiv ID:
arXiv:2609.02451[cs.LG]- Subjects: Machine Learning (
cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)- Submitted on: 2 September 2026
- Authors: Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov
- DOI: 10.48550/arXiv.2609.02451
- arXiv ID:
arXiv:2609.02451[cs.LG] - 研究领域: 机器学习 (
cs.LG);人工智能 (cs.AI);计算与语言 (cs.CL) - 提交时间: 2026年9月2日
- 作者: Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov
- DOI: 10.48550/arXiv.2609.02451
🔍 摘要详述 (Abstract)
In this paper, we propose a scalable Kronecker-based approximation that captures cross-layer interactions without storing the entire Fisher matrix, enabling practical Hessian analysis for billion-parameter networks where full computation is infeasible. Our approach reveals consistent vulnerability patterns: value projection layers exhibit the highest sensitivity and strongest cross-layer correlations across multiple model families, while other components exhibit architecture-specific behaviors.
在本文中,我们提出了一种可扩展的基于克罗内克积的近似方法,它能够在不存储整个费雪矩阵的情况下捕获跨层交互,从而为无法进行完整计算的十亿参数网络实现实用的黑塞矩阵分析。我们的方法揭示了具有一致性的脆弱性模式:在多个模型家族中,值投影层表现出最高的敏感性和最强的跨层相关性,而其他组件则表现出特定于架构的行为。
Through extensive experiments on quantization, sparsification, inter-layer corruption, and post-corruption fine-tuning, we demonstrate that our approximation strongly correlates with both performance degradation and recovery. Our framework provides a practical, theoretically grounded tool for identifying fragile components in large models, opening new avenues for guided compression and optimization strategies, such as mixed-precision allocation, layer-wise sparsity, and adaptive low-rank decomposition across layers and even individual weight groups.
通过在量化、稀疏化、层间损坏以及损坏后微调方面的广泛实验,我们证明了我们的近似方法与性能退化和恢复之间存在高度相关性。我们的框架为识别大型模型中的脆弱组件提供了一个实用且有理论基础的工具,为指导压缩和优化策略(如混合精度分配、分层稀疏性以及跨层甚至单个权重组的自适应低秩分解)开辟了新的途径。
🔗 链接与资源 (Links & Resources)
- Full-Text Access: View PDF | HTML Version | TeX Source
- License: Creative Commons Attribution 4.0
- External Citations: Google Scholar | Semantic Scholar | NASA ADS
- 全文访问: 查看 PDF | HTML 版本 | TeX 源码
- 开源协议: 知识共享署名 4.0

- 外部引用: Google 学术 | Semantic Scholar | NASA ADS