跳转至

SignVerse-2M:一个包含 55 种以上手语、拥有两百万剪辑的姿态原生宇宙

文章背景与核心概要

随着计算机视觉和多模态人工智能的发展,手语识别、翻译及视频生成技术受到了广泛关注。然而,传统的大规模手语数据集多在受控的实验室环境中录制,往往难以泛化到开放世界场景,且无法很好地与现代基于姿态驱动的生成框架(如使用 DWPose 的框架)相结合。

为了弥补这一空白,本文推出了 SignVerse-2M,这是一个大规模的多语言手语数据集和资源库,包含约两百万个视频剪辑,涵盖 55 种以上的手语。该研究通过将 DWPose 作为标准化的预处理流水线,去除了背景和外观变化(如服装或光照条件),同时保留了真实世界的说话者多样性和自然录制条件。此外,作者还引入了数据构建流水线、明确的任务定义以及一个名为 SignDW Transformer 的基线模型,充分展示了该数据集在现代姿态空间建模中的有效性和兼容性。

SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages

Summary

SignVerse-2M is a large-scale multilingual sign language dataset and resource designed for pose modeling, recognition, translation, and video generation. While traditional large-scale sign language datasets provide raw video-text alignments recorded mostly in controlled laboratory environments, they often fail to generalize to open-world scenarios or integrate with modern pose-driven generation frameworks (such as those using DWPose).

To bridge this gap, SignVerse-2M converts publicly available multilingual sign language videos into a unified corpus of approximately two million clips spanning over 55 sign languages. By using DWPose as a standardized preprocessing pipeline, the dataset strips away background and appearance variations (like clothing or lighting conditions) while preserving real-world speaker diversity and natural recording conditions. The authors also introduce a data construction pipeline, clear task definitions, and a baseline model called SignDW Transformer to demonstrate the dataset's efficacy and compatibility with modern pose-space modeling.


Metadata & Reference Information


现有的大规模手语资源通常仅在原始视频与文本对齐的层面上提供监督,且大多在实验室环境中制作。虽然这些资源对于语义理解很重要,但它们无法直接为开放世界的识别和翻译,或者为现代基于姿态驱动的手语视频生成框架提供一个统一的接口:

  1. 基于 RGB 的预训练识别模型在录制过程中严重依赖固定的背景或服装条件,在开放世界设置中的鲁棒性不如风格无关的姿态处理模型。
  2. 近期基于姿态引导的图像/视频生成模型大多使用统一的关键点表示(如 DWPose)作为其控制接口。目前,手语领域仍然缺乏能够直接与这种现代姿态原生范式对接,同时又面向真实世界开放场景的数据资源。

我们提出了 SignVerse-2M,这是一个用于手语姿态建模和评估的大规模多语言姿态原生数据集。它构建于公开可用的多语言手语视频资源之上,通过统一的预处理流水线应用 DWPose,将原始视频转换为可直接用于建模的 2D 姿态序列,从而形成了一个包含约两百万个剪辑、覆盖 55 种以上手语的整合语料库。

与许多实验室数据集不同,该资源保留了真实世界视频的录制条件和说话者多样性,同时通过统一的姿态表示减少了外观变化。为此,我们进一步提供了数据构建流水线、任务定义以及一个简单的 SignDW Transformer 基线模型,证明了该资源在多语言姿态空间建模中的可行性及其与现代姿态驱动流水线的兼容性,同时讨论了它所支持的评估声明及其当前的局限性。

Abstract

Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings. While such resources are important for semantic understanding, they do not directly provide a unified interface for open-world recognition and translation, or for modern pose-driven sign language video generation frameworks:

  1. RGB-based pretrained recognition models depend heavily on fixed backgrounds or clothing conditions during recording, and are less robust in open-world settings than style-agnostic pose-processing models.
  2. Recent pose-guided image/video generation models mostly use a unified keypoint representation such as DWPose as their control interface. At present, the sign language field still lacks a data resource that can directly interface with this modern pose-native paradigm while also targeting real-world open scenarios.

We present SignVerse-2M, a large-scale multilingual pose-native dataset for sign language pose modeling and evaluation. Built from publicly available multilingual sign language video resources, it applies DWPose in a unified preprocessing pipeline to convert raw videos into 2D pose sequences that can be used directly for modeling, resulting in a consolidated corpus of about two million clips covering more than 55 sign languages.

Unlike many laboratory datasets, this resource preserves the recording conditions and speaker diversity of real-world videos while reducing appearance variation through a unified pose representation. Toward this goal, we further provide the data construction pipeline, task definitions, and a simple SignDW Transformer baseline, demonstrating the feasibility of this resource for multilingual pose-space modeling and its compatibility with modern pose-driven pipelines, while discussing the evaluation claims it can support as well as its current limitations.


快速链接与资源: