文章背景与核心概要
高速公路视频异常检测对于维护交通安全至关重要,然而在识别表现出微妙异常运动的远距离车辆时,这仍然是一项巨大的挑战。尽管视觉语言模型(VLM)提供了强大的语义推理能力,但处理连续的全帧视频会稀释来自远端目标的证据,并引入过高的计算开销。
为了克服这些局限性,作者引入了 VIBES 异步框架,该框架利用贝叶斯推理来指导聚焦式的 VLM 推理。在线运动引导的贝叶斯推理模块从车辆轨迹中持续估计与上下文相关的正常运动分布,并更新概率边界;偏离边界则会触发异步警报,从而在时间和空间上隔离出候选异常。此外,VLM 无需处理沉重的连续全帧视频流,而是仅评估与活跃触发器相关的选定帧和局部视觉区域。
实验结果表明,在各种高速公路环境下,VIBES 在保持实时处理效率的同时,显著提高了远距离异常检测的准确性和语义理解能力。
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
📌 Summary
📌 Summary
Expressway video anomaly detection is critical for maintaining traffic safety, yet it remains a formidable challenge—particularly when identifying far-field vehicles exhibiting subtle abnormal motions. While Vision-Language Models (VLMs) offer powerful semantic reasoning capabilities, processing continuous full-frame video dilutes evidence from distant targets and introduces prohibitive computational overhead.
Expressway video anomaly detection is critical for maintaining traffic safety, yet it remains a formidable challenge—particularly when identifying far-field vehicles exhibiting subtle abnormal motions. While Vision-Language Models (VLMs) offer powerful semantic reasoning capabilities, processing continuous full-frame video dilutes evidence from distant targets and introduces prohibitive computational overhead.
To overcome these limitations, the authors introduce VIBES, an asynchronous framework that leverages Bayesian inference to guide focused VLM reasoning. * Motion Estimation & Triggers: An online kinematics-guided Bayesian inference module continuously estimates context-dependent normal-motion distributions from vehicle trajectories, updating probabilistic boundaries. Deviations trigger asynchronous alerts that isolate candidate anomalies in both time and space. * Targeted VLM Reasoning: Rather than processing heavy, continuous full-frame video feeds, the VLM evaluates only selected frames and localized visual regions associated with active triggers.
To overcome these limitations, the authors introduce VIBES, an asynchronous framework that leverages Bayesian inference to guide focused VLM reasoning. * Motion Estimation & Triggers: An online kinematics-guided Bayesian inference module continuously estimates context-dependent normal-motion distributions from vehicle trajectories, updating probabilistic boundaries. Deviations trigger asynchronous alerts that isolate candidate anomalies in both time and space. * Targeted VLM Reasoning: Rather than processing heavy, continuous full-frame video feeds, the VLM evaluates only selected frames and localized visual regions associated with active triggers.
Experimental results demonstrate that VIBES significantly improves far-field anomaly detection accuracy and semantic interpretation while maintaining real-time processing efficiency across varied expressway conditions.
Experimental results demonstrate that VIBES significantly improves far-field anomaly detection accuracy and semantic interpretation while maintaining real-time processing efficiency across varied expressway conditions.
📋 Paper Metadata
📋 Paper Metadata
| Field | Details |
|---|---|
| arXiv Identifier | arXiv:2604.23724 [cs.CV] |
| Primary Subject | Computer Vision and Pattern Recognition (cs.CV) |
| Secondary Subjects | Artificial Intelligence (cs.AI) |
| Authors | Xiaowei Mao, Bowen Sui, Weijie Zhang, Yawen Yang, Shengnan Guo, Shilong Zhao, Jiaqi Lin, Tingrui Wu, Youfang Lin, Huaiyu Wan |
| Submission History | • v1: 26 Apr 2026 • v4 (Latest): 12 Aug 2026 |
| License | Creative Commons Attribution 4.0 |
Field Details arXiv Identifier arXiv:2604.23724 [cs.CV] Primary Subject Computer Vision and Pattern Recognition ( cs.CV)Secondary Subjects Artificial Intelligence ( cs.AI)Authors Xiaowei Mao, Bowen Sui, Weijie Zhang, Yawen Yang, Shengnan Guo, Shilong Zhao, Jiaqi Lin, Tingrui Wu, Youfang Lin, Huaiyu Wan Submission History • v1: 26 Apr 2026
• v4 (Latest): 12 Aug 2026License Creative Commons Attribution 4.0
🔗 Access Links
🔗 Access Links
- Full-Text Options: View PDF | Experimental HTML | TeX Source
- External Bibliographic Tools: NASA ADS | Google Scholar | Semantic Scholar
- Full-Text Options: View PDF | Experimental HTML | TeX Source
- External Bibliographic Tools: NASA ADS | Google Scholar | Semantic Scholar
(Note: Image links and presentation graphics are preserved according to the document specifications.)
(Note: Image links and presentation graphics are preserved according to the document specifications.)