DF3DV-1K:无干扰新视角合成的大规模数据集与基准
文章背景与核心概要
在计算机视觉与三维重建领域,辐射场(Radiance Fields)和3D高斯泼溅(3D Gaussian Splatting)技术的快速发展实现了照片级真实感的新视角合成。然而,过去针对“无干扰”(distractor-free)新视角合成的研究一直缺乏大规模的真实世界数据集。现有的数据往往无法在同一场景下同时提供干净(clean)和杂乱(cluttered)的图像对,这极大地限制了该方向算法的发展和跨场景泛化能力的评估。
为了填补这一空白,本文推出了 DF3DV-1K——一个包含 1,048 个场景、共 89,924 张图像的大规模真实世界数据集。该数据集由消费级相机拍摄,涵盖了室内外多种环境,跨越了 128 种干扰物类型和 161 个场景主题。此外,作者还精心挑选了包含 41 个场景的子集 DF3DV-41 用于评估方法在极端挑战场景下的鲁棒性,并对 9 种最新的无干扰辐射场方法及 3D 高斯泼溅进行了基准测试。
除了基准测试外,该研究还展示了 DF3DV-1K 的实际应用价值:通过微调基于扩散模型的 2D 增强器,显著提升了辐射场方法在 PSNR 和 LPIPS 指标上的表现。这一工作有望推动无干扰视觉技术的进一步发展,并促进三维重建技术超越特定的场景局限。
摘要 (Summary)
DF3DV-1K 是一个面向无干扰新视角合成(NVS)的大规模真实世界数据集和基准测试套件。为了解决以往缺乏同时包含每个场景干净与杂乱图像的大规模数据集的问题,DF3DV-1K 包含了 1,048 个场景 和 89,924 张图像,这些图像由消费级相机拍摄。它涵盖了 128 种干扰物类型和 161 个场景主题,分布在多样化的室内和室外环境中。
DF3DV-1K is a large-scale real-world dataset and benchmarking suite designed for distractor-free novel view synthesis (NVS). Addressing the previous lack of large-scale datasets with both clean and cluttered images per scene, DF3DV-1K comprises 1,048 scenes and 89,924 images captured using consumer-grade cameras. It spans 128 distractor types and 161 scene themes across diverse indoor and outdoor environments.
该论文引入了 DF3DV-41,这是一个由 41 个场景组成的精选子集,旨在测试方法在具有挑战性的场景中的鲁棒性。作者对九种最近的无干扰辐射场方法以及 3D 高斯泼溅(3D Gaussian Splatting)进行了基准测试,并通过微调基于扩散模型的 2D 增强器展示了该数据集的实际效用,该增强器在 PSNR 和 LPIPS 指标上带来了显著的性能提升。
The paper introduces DF3DV-41, a curated subset of 41 scenes designed to test method robustness in challenging scenarios. The authors benchmarked nine recent distractor-free radiance field methods alongside 3D Gaussian Splatting, and demonstrated the practical utility of the dataset by fine-tuning a diffusion-based 2D enhancer that yields noticeable performance improvements in PSNR and LPIPS metrics.
论文元数据 (Paper Metadata)
- arXiv ID: arXiv:2604.13416 [cs.CV]
- 会议: 已被 ECCV 2026 接收
- 主要学科: 计算机视觉与模式识别 (
cs.CV) - 次要学科: 人工智能 (
cs.AI) - 提交时间线:
- v1: 2026年4月15日
- v4(最新修订版): 2026年8月21日
- arXiv ID: arXiv:2604.13416 [cs.CV]
- Conference: Accepted at ECCV 2026
- Primary Subject: Computer Vision and Pattern Recognition (
cs.CV)- Secondary Subject: Artificial Intelligence (
cs.AI)- Submission Timeline:
- v1: April 15, 2026
- v4 (Latest Revision): August 21, 2026
作者 (Authors)
- Cheng-You Lu
- Yi-Shan Hung
- Wei-Ling Chi
- Hao-Ping Wang
- Charlie Li-Ting Tsai
- Yu-Cheng Chang
- Yu-Lun Liu
- Thomas Do
- Chin-Teng Lin
- Cheng-You Lu
- Yi-Shan Hung
- Wei-Ling Chi
- Hao-Ping Wang
- Charlie Li-Ting Tsai
- Yu-Cheng Chang
- Yu-Lun Liu
- Thomas Do
- Chin-Teng Lin
摘要原文 (Abstract)
辐射场的进展使得照片级真实感的新视角合成成为可能。在几个领域中,已经开发了大规模的真实世界数据集来支持全面的基准测试,并推动超越特定场景重建的进展。然而,对于无干扰辐射场而言,每个场景缺乏同时拥有干净和杂乱图像的大规模数据集,这限制了其发展。
Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been developed to support comprehensive benchmarking and to facilitate progress beyond scene-specific reconstruction. However, for distractor-free radiance fields, a large-scale dataset with clean and cluttered images per scene remains lacking, limiting the development.
为了填补这一空白,我们推出了 DF3DV-1K,这是一个大规模的真实世界数据集,包含 1,048 个场景,每个场景都提供用于基准测试的干净和杂乱图像集。该数据集总共包含 89,924 张使用消费级相机拍摄的图像,以模仿随意的拍摄(casual capture),横跨室内和室外环境中的 128 种干扰物类型和 161 个场景主题。
To address this gap, we introduce DF3DV-1K, a large-scale real-world dataset comprising 1,048 scenes, each providing clean and cluttered image sets for benchmarking. In total, the dataset contains 89,924 images captured using consumer cameras to mimic casual capture, spanning 128 distractor types and 161 scene themes across indoor and outdoor environments.
我们系统地设计了一个包含 41 个场景的精选子集 DF3DV-41,用于评估无干扰辐射场方法在具有挑战性的场景下的鲁棒性。利用 DF3DV-1K,我们对九种最新的无干扰辐射场方法和 3D 高斯泼溅进行了基准测试,识别出了最具鲁棒性的方法和最具挑战性的场景。
A curated subset of 41 scenes, DF3DV-41, is systematically designed to evaluate the robustness of distractor-free radiance field methods under challenging scenarios. Using DF3DV-1K, we benchmark nine recent distractor-free radiance field methods and 3D Gaussian Splatting, identifying the most robust methods and the most challenging scenarios.
除了基准测试,我们还展示了 DF3DV-1K 的一项应用:通过微调基于扩散模型的 2D 增强器来改进辐射场方法,在保留集(例如 DF3DV-41)和 On-the-go 数据集上实现了平均 0.96 dB PSNR 和 0.057 LPIPS 的性能提升。我们希望 DF3DV-1K 能够促进无干扰视觉的发展,并推动超越特定场景方法的进展。
Beyond benchmarking, we demonstrate an application of DF3DV-1K by fine-tuning a diffusion-based 2D enhancer to improve radiance field methods, achieving average improvements of 0.96 dB PSNR and 0.057 LPIPS on the held-out set (e.g., DF3DV-41) and the On-the-go dataset. We hope DF3DV-1K facilitates the development of distractor-free vision and promotes progress beyond scene-specific approaches.
资源与链接 (Resources & Links)
- 项目网站与排行榜: 官方 DF3DV-1K 网站
- 全文获取:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 数据集许可证: 知识共享署名 4.0 国际许可协议 (CC BY 4.0)

- Project Website & Leaderboard: Official DF3DV-1K Website
- Full-Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- Dataset License: Creative Commons Attribution 4.0 International (CC BY 4.0)
引用 (Citations)
BibTeX 引用
@misc{lu2026df3dv1klargescaledatasetbenchmark,
title={DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis},
author={Cheng-You Lu and Yi-Shan Hung and Wei-Ling Chi and Hao-Ping Wang and Charlie Li-Ting Tsai and Yu-Cheng Chang and Yu-Lun Liu and Thomas Do and Chin-Teng Lin},
year={2026},
eprint={2604.13416},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.13416}
}
外部平台
External Platforms