文章背景与核心概要
随着对大规模语音数据的依赖日益增加,隐私保护变得至关重要,然而现有的匿名化方法往往会破坏声学连续性或降低声音多样性,从而损害数据实用性。这会负面影响自动语音识别(ASR)、文本转语音(TTS)和语音情感识别(SER)等下游任务。此外,当前的评估方法通常仅依赖于使用预训练模型的直接测试,无法提供全面的实用性图景。
为了解决这些挑战,作者提出了一种新颖的两阶段框架,在保护语言内容和声学身份的同时维持整体可用性:1. 内容隐私:使用生成式语音编辑模型无缝替换个人身份信息(PII)。2. 声音隐私:引入 F3-VA,这是一个基于流匹配(flow-matching)的匿名化框架,采用三阶段设计以产生多样且独特的匿名说话人。
为了确保全面评估,作者同时使用基于声学和内容的说话人验证指标来评估隐私,并通过完全从头开始训练 ASR、TTS 和 SER 模型来衡量实用性。实验结果表明,与 VoicePrivacy 挑战赛的基线相比,该框架在实现卓越隐私保护的同时将实用性退化降至最低,并在隐私约束下更真实地反映了数据的实用性。
Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
Summary
Summary
The growing reliance on large-scale speech data makes privacy protection critical, yet existing anonymization methods often degrade data utility by disrupting acoustic continuity or reducing vocal diversity. This harms downstream tasks such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech Emotion Recognition (SER). Furthermore, current evaluation practices typically rely solely on direct testing with pretrained models, offering an incomplete picture of utility.
The growing reliance on large-scale speech data makes privacy protection critical, yet existing anonymization methods often degrade data utility by disrupting acoustic continuity or reducing vocal diversity. This harms downstream tasks such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech Emotion Recognition (SER). Furthermore, current evaluation practices typically rely solely on direct testing with pretrained models, offering an incomplete picture of utility.
To resolve these challenges, the authors propose a novel two-stage framework that protects linguistic content and acoustic identity while maintaining overall usability: 1. Content Privacy: Uses a generative speech editing model to seamlessly replace personally identifiable information (PII). 2. Voice Privacy: Introduces F3-VA, a flow-matching-based anonymization framework with a three-stage design to produce diverse and distinct anonymized speakers.
To resolve these challenges, the authors propose a novel two-stage framework that protects linguistic content and acoustic identity while maintaining overall usability: 1. Content Privacy: Uses a generative speech editing model to seamlessly replace personally identifiable information (PII). 2. Voice Privacy: Introduces F3-VA, a flow-matching-based anonymization framework with a three-stage design to produce diverse and distinct anonymized speakers.
To ensure comprehensive assessment, the authors evaluate privacy using both acoustic- and content-based speaker verification metrics, and measure utility by training ASR, TTS, and SER models completely from scratch. Experimental results demonstrate that this framework achieves superior privacy protection with minimal utility degradation compared to VoicePrivacy Challenge baselines, while offering a more realistic reflection of utility under privacy constraints.
To ensure comprehensive assessment, the authors evaluate privacy using both acoustic- and content-based speaker verification metrics, and measure utility by training ASR, TTS, and SER models completely from scratch. Experimental results demonstrate that this framework achieves superior privacy protection with minimal utility degradation compared to VoicePrivacy Challenge baselines, while offering a more realistic reflection of utility under privacy constraints.
Document Metadata
Document Metadata
- arXiv ID: arXiv:2604.17000 [eess.AS]
- Subjects: Audio and Speech Processing (
eess.AS); Artificial Intelligence (cs.AI) - Submitted on: 18 April 2026
- DOI: 10.48550/arXiv.2604.17000
- arXiv ID: arXiv:2604.17000 [eess.AS]
- Subjects: Audio and Speech Processing (
eess.AS); Artificial Intelligence (cs.AI)- Submitted on: 18 April 2026
- DOI: 10.48550/arXiv.2604.17000
Authors
- Yunchong Xiao
- Yuxiang Zhao
- Ziyang Ma
- Shuai Wang
- Kai Yu
- Jiachun Liao
- Xie Chen
Authors
- Yunchong Xiao
- Yuxiang Zhao
- Ziyang Ma
- Shuai Wang
- Kai Yu
- Jiachun Liao
- Xie Chen
Access Links
Access Links
- Full-Text: View PDF | HTML (experimental) | TeX Source
- External Resources: NASA ADS | Google Scholar | Semantic Scholar
- Full-Text: View PDF | HTML (experimental) | TeX Source
- External Resources: NASA ADS | Google Scholar | Semantic Scholar
view license