文章背景与核心概要
近年来,3D高斯溅射(3DGS)技术在实现快速、逼真的说话人脸渲染方面取得了显著进展,但在精确的唇部发音控制上仍面临挑战。传统的连续音频嵌入模型往往会导致口型运动过度平滑,并且容易产生双唇无法闭合等“漏嘴”伪影,无法满足严格的发音约束。
为了解决这一技术瓶颈,本文提出了PD-GS(Phoneme-Driven Gaussian Splatting)框架。该模型通过自动语音识别(ASR)和强制对齐流水线引入了时间对齐的音素标记,并创新性地设计了语言融合模块(LFM),实现了连续音频上下文与显式音素指导的有机平衡。该研究已被国际多媒体顶级会议 ACM MM 2026 接收,在HDTF数据集上的实验证明其在唇部几何精度和减少闭合违规方面具有显著优势。
PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
Executive Summary
PD-GS (Phoneme-Driven Gaussian Splatting) is an advanced framework designed to improve lip articulation and visual fidelity in audio-driven talking head generation.
While traditional 3D Gaussian Splatting (3DGS) models allow for fast, photorealistic rendering, they frequently suffer from over-smoothed mouth movements and "leaky mouth" artifacts (violating hard articulatory constraints like bilabial closures). This happens because continuous acoustic embeddings struggle to reliably infer brief, discrete articulatory events.
To solve this, PD-GS incorporates time-aligned phoneme tokens via an automatic ASR and forced-alignment pipeline, using a novel Linguistic Fusion Module (LFM) to balance continuous audio context with explicit phoneme guidance. Accepted at ACM MM 2026, the model demonstrates superior lip geometry and fewer closure violations on the HDTF dataset.
Paper Overview
Field Details Title PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads Authors Ao Fu, Yi Zhou Submitted August 5, 2026 Venue Accepted to ACM MM 2026 Primary Subject Artificial Intelligence ( cs.AI); Sound (cs.SD)arXiv ID 2608.05218
Abstract
3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious "leaky mouth" artifact.
A key difficulty is that brief, discrete articulatory events are inferred from a continuous acoustic embedding under a regression objective, which biases predictions toward averaged mouth configurations. While modern self-supervised speech encoders provide rich prosodic and phonetic cues, they do not provide an explicit, frame-aligned linguistic target that reliably disambiguates closure-level events.
We propose Phoneme-Driven Gaussian Splatting (PD-GS), which augments a 3DGS talker with time-aligned phoneme tokens obtained from an automatic ASR and forced-alignment pipeline. Our core component, the Linguistic Fusion Module (LFM), adaptively fuses continuous audio context with discrete phoneme embeddings through a learned gate, allowing the model to preserve smooth audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments.
PD-GS is trained purely from monocular video using image reconstruction and lip landmark supervision. On HDTF, PD-GS achieves the best lip geometry among the compared baselines (LMD 2.66) and qualitatively reduces closure violations in challenging phoneme sequences, yielding more linguistically faithful neural avatars.
Key Innovations
- Phoneme-Driven Augmentation: Integrates time-aligned phoneme tokens derived from an ASR and forced-alignment pipeline into a 3DGS talking-head framework.
- Linguistic Fusion Module (LFM): Utilizes a learned gate to adaptively combine continuous acoustic contexts with discrete phoneme embeddings, retaining natural speech dynamics while ensuring precise articulation during critical phonetic segments.
- Robust Supervision: Trained purely from monocular video sources leveraging standard image reconstruction and precise lip landmark supervision.
Results & Performance
- Dataset: Evaluated on the High-Definition Talking Face (HDTF) benchmark.
- Lip Geometry: Achieves state-of-the-art lip geometry among tested baselines with a Landmark Distance (LMD) score of 2.66.
- Visual Quality: Effectively mitigates closure-violation artifacts (such as "leaky mouths") during complex or rapid phoneme sequences.
Full-Text & Access Links
- arXiv Abstract: arXiv:2608.05218
- Direct PDF Download: View PDF
- Experimental HTML: arXiv HTML Version
- License: Creative Commons Attribution 4.0
(Source image assets referenced from original metadata)
