跳转至

文章背景与核心概要

可靠的水下感知长期以来一直受到浑浊环境中光学相机局限性的阻碍。虽然光学传感器能够提供丰富的语义细节,但在能见度差的环境下会迅速退化;相比之下,成像声呐能够提供几何结构,但缺乏语义深度。为此,本文推出了 SonarLLM——一种旨在将声呐作为原生模态的新型多模态大语言模型。通过整合声呐专用编码器、物理感知特征增强以及可靠性感知融合,该模型有效地将声学结构与光学语义进行了对齐。

研究人员还推出了 SonarBench,这是一个用于评估识别、计数和视觉问答(VQA)任务的综合评测基准。实验证明,SonarLLM 显著优于现有基线模型,特别是在光学条件恶化的情况下,其优势更为明显。


SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

Authors: Cong Su, Longxuan Ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
Date: August 25, 2026
arXiv ID: 2608.24325
Subject: Artificial Intelligence (cs.AI)


Summary

Reliable underwater perception is hindered by the limitations of optical cameras in turbid environments. While optical sensors provide rich semantic detail, they degrade rapidly in poor visibility, whereas imaging sonar offers geometric structure but lacks semantic depth. SonarLLM is a novel multimodal large language model designed to treat sonar as a native modality. By integrating a sonar-specific encoder, physics-aware feature enhancement, and reliability-aware fusion, the model effectively aligns acoustic structures with optical semantics. The researchers also introduce SonarBench, a comprehensive benchmark for evaluating recognition, counting, and VQA tasks, demonstrating that SonarLLM significantly outperforms existing baselines, particularly as optical conditions deteriorate.

可靠的水下感知长期以来一直受到浑浊环境中光学相机局限性的阻碍。虽然光学传感器能够提供丰富的语义细节,但在能见度差的环境下会迅速退化;相比之下,成像声呐能够提供几何结构,但缺乏语义深度。SonarLLM 是一种新型多模态大语言模型,旨在将声呐作为原生模态进行处理。通过集成声呐专用编码器、物理感知特征增强以及可靠性感知融合,该模型有效地将声学结构与光学语义对齐。研究人员还引入了 SonarBench,这是一个用于评估识别、计数和 VQA 任务的综合评测基准,证明了 SonarLLM 显著优于现有的基线模型,特别是在光学条件恶化的情况下。


Key Contributions

  • Native Sonar Integration: Unlike existing MLLMs that rely primarily on optical encoders, SonarLLM incorporates a dedicated sonar encoder to process acoustic data as a primary input modality.
  • Physics-Aware Feature Enhancement: The model utilizes modality-specific processing to handle the unique range-azimuth structure and acoustic artifacts inherent in sonar imaging.
  • Reliability-Aware Hierarchical Fusion: A dynamic fusion mechanism that adjusts the contribution of sonar and optical inputs based on the sensing quality, ensuring robust performance under varying turbidity.
  • SonarBench Benchmark: A new paired dataset spanning four tasks (recognition, counting, visual question answering, and captioning) across three input settings (sonar-only, optical-only, and fusion) to measure cross-modal complementarity.

核心贡献

  • 原生声呐集成: 与主要依赖光学编码器的现有 MLLM 不同,SonarLLM 整合了一个专用的声呐编码器,将声学数据作为主要的输入模态进行处理。
  • 物理感知特征增强: 该模型利用模态特定的处理方法来应对声呐成像中固有的独特距离-方位角结构以及声学伪影。
  • 可靠性感知分层融合: 采用动态融合机制,根据传感质量调整声呐和光学输入的贡献,从而确保在不同浑浊度下具有稳健的性能。
  • SonarBench 评测基准: 一个包含四个任务(识别、计数、视觉问答和图像描述)以及三种输入设置(仅声呐、仅光学和融合)的新型配对数据集,用于衡量跨模态的互补性。

Performance Highlights

  • Sonar-Only Performance: Achieved 72.0% macro accuracy across recognition, counting, and VQA tasks, outperforming the strongest baseline by 34.4 percentage points.
  • Fusion Performance: Achieved 68.7% accuracy, exceeding the best baseline by 25.1 points.
  • Adaptive Robustness: In scenarios with increasing turbidity, the fusion-over-optical gain for recognition and counting tasks grew from 6.0 to 36.0 points, proving the model's ability to effectively leverage sonar as a fallback and complementary sensor.

性能亮点

  • 纯声呐性能: 在识别、计数和 VQA 任务中实现了 72.0% 的宏平均准确率,比最强的基线模型高出 34.4 个百分点
  • 融合性能: 实现了 68.7% 的准确率,超出了最佳基线模型 25.1 个百分点
  • 自适应鲁棒性: 在浑浊度不断增加的场景中,融合模型相对于纯光学模型在识别和计数任务上的增益从 6.0 个百分点增长到了 36.0 个百分点,证明了该模型能够有效利用声呐作为备用和互补传感器的能力。

Access & Resources

访问与资源