跳转至

多模态大语言模型不断演进的安全格局:新兴威胁与防御综述

文章背景与核心概要

随着多模态大语言模型(MLLMs)将文本与图像、音频、视频等异构模态深度融合,其理解与推理能力得到了显著增强。然而,这种跨模态的架构转变也彻底改变了机器学习的安全范式。系统复杂性的增加以及多模态交互带来了独特的安全漏洞(如模态集成受损、模态错位以及复合安全风险),这些都是传统的单模态安全框架所无法应对的。

本文综述对 MLLM 不断演变的安全格局进行了全面分析。文章提出了一种基于多模态落地的威胁分类体系(涵盖对抗攻击、数据 poisoning、越狱以及幻觉),深入审视了更新后的安全假设,系统总结了最新的防御机制,并勾勒出构建具可扩展性且合乎规范的可信 AI 所面临的开放性挑战与未来方向。


📌 Summary

Multi-modal Large Language Models (MLLMs) bridge text and heterogeneous modalities like images, audio, and video to enhance comprehension and reasoning capabilities. However, this cross-modal integration significantly alters the machine learning security paradigm. Increased system complexity and multi-modal interactions introduce unique vulnerabilities—such as compromised modality integration, modality misalignment, and compounded safety risks—that traditional, uni-modal security frameworks cannot adequately address.

This survey provides a comprehensive analysis of the shifting security landscape for MLLMs, introducing a multi-modal-grounded taxonomy of threats (including adversarial attacks, data poisoning, jailbreaks, and hallucinations), examining updated safety assumptions, synthesizing recent defense mechanisms, and outlining open challenges for building scalable, principled trustworthy AI.


📋 Document Information

Metadata Field Details
arXiv Identifier arXiv:2608.07535 [cs.LG]
Primary Subject Machine Learning (cs.LG)
Secondary Subjects Artificial Intelligence (cs.AI), Computers and Society (cs.CY)
Submission Date July 27, 2026
Venue Accepted at the ICLR 2026 Workshop on Principled Design for Trustworthy AI
DOI 10.48550/arXiv.2608.07535

👥 Authors

  • Xi Li
  • Shu Zhao
  • Xiaohan Zou
  • Fei Zhao
  • Fuxiao Liu
  • Yusen Zhang
  • Cheng Han
  • Yushun Dong
  • Jiaqi Wang

📖 Abstract

Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of machine learning. Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fused safety risks, reflecting shifts in threat modeling beyond uni-modal assumptions. These shifts, in turn, impose new constraints on safety solutions not captured by existing frameworks rooted in uni-modal learning. Motivated by these challenges, this survey provides a systematic analysis of the evolving safety landscape of MLLMs. We first propose a multimodal grounded taxonomy of safety threats and analyze shifts in threat models, covering adversarial attacks, data poisoning, jailbreaks, and hallucinations. We then summarize updated safety assumptions and organize recent advances in MLLM safety strategies accordingly. Finally, we discuss open challenges and future directions to inform the development of more principled and scalable safety mechanisms for multimodal systems.