倾听潜变量:通过大音频语言模型中的隐状态交互实现自纠错语音识别
文章背景与核心概要
近年来,自动语音识别(ASR)技术的发展频繁引入大语言模型(LLMs),以期通过语义理解来提升转录的准确率。尽管诸如对数几率融合(外部)和热启动初始化(内部)等策略在业界很常见,但如何将它们进行高效结合仍然是一个亟待解决的挑战。
本文引入了 Hybrid Search(混合搜索),这是一个针对 LoRA 微调设置下热启动 LLM-based ASR 模型的新型自纠错框架。该研究的核心见解与贡献包括:1. 隐状态交互信号:将基于 LLM 的 ASR 隐状态与微调前基础 LLM 的隐状态相链接的交互特征,能够为标记(token)的语义依赖性提供强有力的信号;2. 定向精炼:与依赖幼稚的全局纠错方法(如重打分或后期融合)不同,有选择性地精炼高语义依赖性的标记可大幅提升 ASR 性能;3. 推理时增强:研究结果表明,即便在通过热启动完成知识迁移之后,模型在推理阶段依然可以利用其基础 LLM 来进一步提升性能。
1. Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
arXiv ID: arXiv:2609.02940 [cs.CL]
Accepted to: Findings of EMNLP 2026
Submission Date: August 31, 2026
arXiv ID: arXiv:2609.02940 [cs.CL]
Accepted to: Findings of EMNLP 2026
Submission Date: August 31, 2026
Authors
- Chan-Jan Hsu
- Jaeyeon Kim
- Chao-Han Huck Yang
- Shinji Watanabe
- Hung-yi Lee
- Carlos Busso
Authors
- Chan-Jan Hsu
- Jaeyeon Kim
- Chao-Han Huck Yang
- Shinji Watanabe
- Hung-yi Lee
- Carlos Busso
Summary
Recent advancements in Automatic Speech Recognition (ASR) frequently incorporate Large Language Models (LLMs) to enhance transcription accuracy through semantic understanding. While strategies like logit fusion (external) and warm initialization (internal) are common, combining them effectively remains an open challenge.
This paper introduces Hybrid Search, a novel self-correction framework for warm-initialized LLM-based ASR models in LoRA-adapted settings. The core insights and contributions of the work include: 1. Hidden-State Interaction Signals: Interaction features linking the hidden states of the LLM-based ASR and the pre-adaptation base LLM provide strong signals regarding a token's semantic dependence. 2. Targeted Refinement: Instead of relying on naive global correction methods (such as rescoring or late fusion), selectively refining tokens with high semantic dependence substantially improves ASR performance. 3. Inference-Time Enhancement: The findings demonstrate that even after knowledge is transferred via warm initialization, models can still leverage their base LLM to boost performance during inference.
Summary
Recent advancements in Automatic Speech Recognition (ASR) frequently incorporate Large Language Models (LLMs) to enhance transcription accuracy through semantic understanding. While strategies like logit fusion (external) and warm initialization (internal) are common, combining them effectively remains an open challenge.
This paper introduces Hybrid Search, a novel self-correction framework for warm-initialized LLM-based ASR models in LoRA-adapted settings. The core insights and contributions of the work include: 1. Hidden-State Interaction Signals: Interaction features linking the hidden states of the LLM-based ASR and the pre-adaptation base LLM provide strong signals regarding a token's semantic dependence. 2. Targeted Refinement: Instead of relying on naive global correction methods (such as rescoring or late fusion), selectively refining tokens with high semantic dependence substantially improves ASR performance. 3. Inference-Time Enhancement: The findings demonstrate that even after knowledge is transferred via warm initialization, models can still leverage their base LLM to boost performance during inference.
Access Links & Resources
- Paper & Documentation:
- View PDF
- arXiv:2609.02940 HTML (Experimental)
- TeX Source
- DOI (DataCite)
- License: Creative Commons Attribution 4.0
(View License) - External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS
Access Links & Resources
- Paper & Documentation:
- View PDF
- arXiv:2609.02940 HTML (Experimental)
- TeX Source
- DOI (DataCite)
- License: Creative Commons Attribution 4.0
(View License)
- External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS