跳转至

文章背景与核心概要

生成式人工智能的飞速发展使得高度逼真的深度伪造(Deepfake)视频层出不穷,对数字信任和AI安全构成了严重威胁。然而,现有的深度伪造数据集往往缺乏多样化、最先进的生成方法以及可靠的细粒度标注,而传统的检测器和多模态大语言模型(MLLMs)在捕捉不同场景下的细微伪造痕迹时往往显得力不从心。

为了克服这些挑战,本文推出了大规模基准测试 FaceVid-Forensics-100K,其中包含100,000个视频,涵盖33种合成方法,并配备了通过MLLM驱动的流水线自动生成的细粒度文本标注。此外,本文还提出了一种多智能体取证推理框架,通过四个专注于纹理光照运动物理的领域专家智能体从不同角度分析视频,最后由裁判智能体综合其分析结果。尽管该系统仅使用小型开源MLLM,但其性能超越了闭源模型(如GPT和Gemini),并在该基准测试中达到了最先进的水平。


Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

arXiv ID: arXiv:2608.06865 [cs.CV]
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
Submitted: August 7, 2026
Authors: Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing
Project Page: ARGUS Project Page


📌 Summary

摘要概览:生成式AI的迅猛发展使得逼真深度伪造视频的创建对数字信任与安全构成了严重威胁。现有的深度伪造数据集往往缺乏多样化、最先进的生成方法以及可靠的细粒度标注,同时传统检测器与多模态大语言模型(MLLMs)在各种场景下捕捉细微伪造痕迹时表现不佳。

为了解决这些问题,本文引入了: 1. FaceVid-Forensics-100K:一个包含100,000个视频的大规模基准数据集,跨越33种合成方法(包括Seedance 2.0等现代生成器),具备通过MLLM驱动的流水线自动生成的细粒度文本标注与解释。 2. 多智能体取证推理框架:该系统利用四个专门的领域专家智能体,从不同角度——纹理(texture)光照(lighting)运动(motion)物理(physics)——分析视频,随后由一个裁判智能体整合它们的发现。尽管仅使用小型开源MLLM,该方法仍超越了闭源模型(如GPT和Gemini),并在该基准上取得了最先进的性能。

The rapid advancement of generative AI has made the creation of realistic deepfake videos a severe threat to digital trust and security. Existing deepfake datasets often lack diverse, state-of-the-art generation methods and reliable fine-grained annotations, while conventional detectors and Multimodal Large Language Models (MLLMs) struggle to catch subtle forgery artifacts across varied scenarios.

To overcome these issues, this paper introduces: 1. FaceVid-Forensics-100K: A large-scale benchmark of 100,000 videos spanning 33 synthesis methods (including modern generators like Seedance 2.0), featuring fine-grained textual annotations and explanations automatically generated via an MLLM-powered pipeline. 2. A Multi-Agent Forensic Reasoning Framework: A system leveraging four specialized domain-expert agents to analyze videos from distinct angles—texture, lighting, motion, and physics—followed by a judge agent that consolidates their findings. Despite using only small open-source MLLMs, this approach outperforms closed-source models (like GPT and Gemini) and achieves state-of-the-art performance on the benchmark.


👥 Authors

作者团队: * Xuechao Zou * Shun Zhang * Kai Li * Yi Zhou * Xinyu Sun * Yuhui Chen * Zhe Wu * Congyan Lang * Junliang Xing

  • Xuechao Zou
  • Shun Zhang
  • Kai Li
  • Yi Zhou
  • Xinyu Sun
  • Yuhui Chen
  • Zhe Wu
  • Congyan Lang
  • Junliang Xing

📄 Abstract

论文摘要:利用生成式人工智能恶意创建高度逼真的深度伪造视频引发了严重的伦理担忧,并对人工智能安全构成了实质性挑战。然而,现有的深度伪造视频基准对近期合成方法的覆盖有限,且普遍缺乏可靠的细粒度文本标注。同时,传统的检测器和多模态大语言模型(MLLMs),无论是作为单一模型运行还是依赖单一分析视角,往往无法捕捉细微的伪造痕迹,从而限制了它们对新兴AI生成方法的泛化能力。

为了解决这些局限性,我们引入了 FaceVid-Forensics-100K,这是一个大规模的深度伪造视频数据集,包含100,000个视频,涵盖换脸、面部重演和全脸合成等33种合成方法,其中包括Seedance 2.0等近期生成器。该数据集提供了视觉观测的细粒度文本标注以及与判定一致的取证解释,这些内容是通过由先进MLLM驱动的多模型聚合和冲突解决流水线自动合成的。

基于此基准,我们提出了一种多智能体取证推理框架,该框架采用四个专门的领域专家智能体,从四个视角独立分析伪造线索:纹理光照运动物理。随后,裁判智能体协调它们的报告,生成最终预测及解释。在域外测试集上的广泛评估表明,尽管我们的框架完全由小型开源MLLM组成,但其性能超越了所有方法(包括闭源的GPT和Gemini模型),并在该基准报告的所有指标中名列第一。

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI-generated methods.

To address these limitations, we introduce FaceVid-Forensics-100K, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine-grained textual annotations of visual observations and verdict-consistent forensic explanations, automatically synthesized through a multi-model aggregation and conflict-resolution pipeline powered by advanced MLLMs.

Building on this benchmark, we propose a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out-of-domain test sets show that, despite being composed entirely of small open-source MLLMs, our framework outperforms all methods including closed-source GPT and Gemini models and ranks first across all reported metrics on this benchmark.


📊 Document Metadata & Resources

文档元数据与资源: * 论文长度: 22页,8幅图,14张表 * 全文链接: * 查看 PDF * HTML 版本(实验性) * TeX 源码 * 引用与参考: * NASA ADS * Google Scholar * Semantic Scholar * DOI (DataCite)