跳转至

文章背景与核心概要

在真实世界中,音频信号往往会遭受复杂的失真耦合(Complex Distortion Coupling),同时用户对音频处理也有着高度的个性化需求。传统的音频增强方案往往难以同时兼顾这两方面的挑战。为此,本文推出了 StrixAE——一种基于多模态大语言模型(MLLM)的创新智能体,旨在通过将 MLLM 作为中央控制器,无缝协调多个音频增强与个性化模型,从而突破传统方法的局限。

为了实现卓越的鲁棒性、伪影消除以及跨场景泛化能力,该智能体采用了严谨的两阶段训练流程:首先在 AcoustBench 上进行思维链(CoT)监督微调,建立基础的推理和工具调用能力;随后引入音频感知强化学习(APRL),这是一种专门为音频恢复流水线量身定制的奖励机制,能够同时优化格式有效性、结构连贯性和感知质量。与标准的强化学习微调不同,APRL 结合了结构化奖励,以保证生成可执行的流水线并确保正确的逻辑分段顺序,从而彻底消除了工具幻觉现象。在真实世界数据集上的广泛评估表明,StrixAE 的性能超越了领先的开源和闭源替代方案,在各项感知指标上树立了全新的技术先进水平(SOTA)。


StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios

Summary

StrixAE is an innovative multimodal large language model (MLLM)-based intelligent agent designed to tackle real-world audio enhancement challenges characterized by complex distortion couplings and the need for personalized processing. Traditional solutions often struggle to handle both aspects simultaneously. StrixAE overcomes this limitation by employing an MLLM as a central controller to seamlessly coordinate multiple audio enhancement and personalization models.

To achieve superior robustness, artifact reduction, and cross-scenario generalization, the agent is trained through a rigorous two-stage process: 1. Chain-of-Thought (CoT) Supervised Fine-Tuning on AcoustBench to establish foundational reasoning and tool-invocation capabilities. 2. Audio Perception Reinforcement Learning (APRL), a reward mechanism tailored for audio restoration pipelines that optimizes format validity, structural coherence, and perceptual quality simultaneously.

Unlike standard reinforcement learning fine-tuning, APRL incorporates structured rewards to guarantee executable pipelines and proper logical section ordering, eliminating tool hallucinations. Extensive evaluations on real-world datasets demonstrate that StrixAE outperforms leading open-source and proprietary alternatives, setting a new state-of-the-art across various perceptual metrics.


Document Metadata

Field Details
arXiv ID arXiv:2609.03414
Primary Subject Sound (cs.SD); Artificial Intelligence (cs.AI)
Submission Date September 3, 2026
Authors Chenglin Wu, Junjie Wu, Jinhang Chen, Mingyang Chen, Zixu Lin, Jiabian Chen, Xinghao Ding, Xiaotong Tu