文章背景与核心概要
儿童语音中的音素识别长期以来一直受到训练数据匮乏以及年轻说话者独特声学特征的限制。本文介绍了 PhonemeTrainer,这是一个轻量级、兼容边缘端且能够在现代智能手机上运行的模型。通过引入年龄感知训练(即模型在学习音素序列的同时预测说话者的年龄),作者使一个 94M 参数的模型实现了高性能。
这种方法超越了显著更大的模型(例如 317M 参数的 WavLM Large),并取得了与参数量达到其 90 倍的大型集成模型相媲美的结果。该技术为儿童语音识别和教育应用提供了一种保护隐私的端侧解决方案,具备极高的实际应用价值。
Edge Phoneme Recognition for Children's Speech through Age-Aware Training
arXiv: 2608.10206
Submitted: August 10, 2026
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)
Summary
Phoneme recognition in children’s speech has historically been hindered by data scarcity and the unique acoustic characteristics of younger speakers. This paper introduces PhonemeTrainer, a lightweight, edge-compatible model capable of running on modern smartphones. By incorporating age-aware training—where the model learns to predict the speaker's age alongside the phoneme sequence—the authors achieved high performance with a 94M-parameter model. This approach outperformed significantly larger models (such as the 317M-parameter WavLM Large) and achieved results comparable to massive ensemble models, offering a privacy-preserving solution for children's speech recognition and educational applications.
Summary
Phoneme recognition in children’s speech has historically been hindered by data scarcity and the unique acoustic characteristics of younger speakers. This paper introduces PhonemeTrainer, a lightweight, edge-compatible model capable of running on modern smartphones. By incorporating age-aware training—where the model learns to predict the speaker's age alongside the phoneme sequence—the authors achieved high performance with a 94M-parameter model. This approach outperformed significantly larger models (such as the 317M-parameter WavLM Large) and achieved results comparable to massive ensemble models, offering a privacy-preserving solution for children's speech recognition and educational applications.
Authors
- Matthew Arboleda
- Ryan Arboleda
- Sophie Haak
- Sam Hjelmeset
- Andrew Franck
- Bingrui Yang
- Jose Bustamante Ortiz
- Yuanrong Shen
- Joel Walsh
Authors
- Matthew Arboleda
- Ryan Arboleda
- Sophie Haak
- Sam Hjelmeset
- Andrew Franck
- Bingrui Yang
- Jose Bustamante Ortiz
- Yuanrong Shen
- Joel Walsh
Abstract
Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.04 CER of competition ensembles with 90 times the parameters. This has enabled the creation of PhonemeTrainer, an application that can run on most modern cellular phones. This will ultimately enable better Automated Speech Recognition (ASR) and pronunciation helper apps for children's speech, with the privacy and compliance benefits that come with edge processing.
Abstract
Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.04 CER of competition ensembles with 90 times the parameters. This has enabled the creation of PhonemeTrainer, an application that can run on most modern cellular phones. This will ultimately enable better Automated Speech Recognition (ASR) and pronunciation helper apps for children's speech, with the privacy and compliance benefits that come with edge processing.
Metadata & Access
- Comments: 3 pages, 2 figures, 1 table. Presented at the 13th ACM Conference on Learning @ Scale (L@S '26).
- DOI: https://doi.org/10.48550/arXiv.2608.10206
- License:
view license
Full-Text Links
Metadata & Access
- Comments: 3 pages, 2 figures, 1 table. Presented at the 13th ACM Conference on Learning @ Scale (L@S '26).
- DOI: https://doi.org/10.48550/arXiv.2608.10206
- License:
view license
Full-Text Links