Personalized ASR Models from a Large and Diverse Disordered Speech Dataset
个性化ASR模型:基于大规模多样化言语障碍语音数据集
Posted by Katrin Tomanek, Software Engineer and Bob MacDonald, Technical Program Manager, Google Research
作者:Katrin Tomanek,软件工程师;Bob MacDonald,Google Research 技术项目经理
Speech impairments affect millions of people, with underlying causes ranging from neurological or genetic conditions to physical impairment, brain damage or hearing loss. Similarly, the resulting speech patterns are diverse, including stuttering, dysarthria, apraxia, etc., and can have a detrimental impact on self-expression, participation in society and access to voice-enabled technologies. Automatic speech recognition (ASR) technologies have the potential to help individuals with such speech impairments by improving access to dictation and home automation and by enhancing communication. However, while the increased computational power of deep learning systems and the availability of large training datasets has improved the accuracy of ASR systems, their performance is still insufficient for many people with speech disorders, rendering the technology unusable for many of the speakers who could benefit the most.
言语障碍影响着数百万人,其根本原因涵盖神经系统或遗传性疾病、身体损伤、脑损伤或听力损失等。同样,由此产生的言语模式也多种多样,包括口吃、构音障碍、言语失用等,并且可能对自我表达、社会参与以及语音技术的使用产生不利影响。自动语音识别(ASR)技术有潜力通过改善听写、家庭自动化以及增强沟通,来帮助有此类言语障碍的个体。然而,尽管深度学习系统计算能力的提升和大规模训练数据集的可用性提高了ASR系统的准确性,但其性能对许多言语障碍者来说仍然不足,使得这项技术对许多最可能受益的使用者来说无法使用。
In 2019, we introduced Project Euphonia and discussed how we could use personalized ASR models of disordered speech to achieve accuracies on par with non-personalized ASR on typical speech. Today we share the results of two studies, presented at Interspeech 2021, that aim to expand the availability of personalized ASR models to more users. In “Disordered Speech Data Collection: Lessons Learned at 1 Million Utterances from Project Euphonia”, we present a greatly expanded collection of disordered speech data, composed of over 1 million utterances. Then, in “Automatic Speech Recognition of Disordered Speech: Personalized models outperforming human listeners on short phrases”, we discuss our efforts to generate personalized ASR models based on this corpus. This approach leads to highly accurate models that can achieve up to 85% improvement to the word error rate (WER) in select domains compared to out-of-the-box speech models trained on typical speech.
2019年,我们推出了Project Euphonia,并讨论了如何使用个性化ASR模型来处理障碍言语,使其准确率能达到与针对典型语音的非个性化ASR相当的水平。今天,我们分享两项研究的结果(发表于Interspeech 2021),旨在将个性化ASR模型的应用扩展到更多用户。在《障碍言语数据收集:Project Euphonia百万条语音数据的经验教训》中,我们展示了一个大幅扩展的障碍言语数据集,包含超过100万条语音。随后,在《障碍言语自动语音识别:个性化模型在短句上优于人类听众》中,我们讨论了基于此语料库构建个性化ASR模型的工作。这种方法产生了高精度的模型,在特定领域相比未适配的典型语音模型,单词错误率(WER)可降低高达85%。
Impaired Speech Data Collection
障碍言语数据收集
Since 2019, speakers with speech impairments of varying degrees of severity across a variety of conditions have provided voice samples to support Project Euphonia’s research mission. This effort has grown Euphonia’s corpus to over 1 million utterances, comprising over 1400 hours from 1330 speakers (as of August 2021).
自2019年以来,来自各种状况、具有不同程度言语障碍的使用者提供了语音样本,以支持Project Euphonia的研究使命。这项工作已将Euphonia的语料库扩展到超过100万条语音,涵盖来自1330位使用者的1400多小时录音(截至2021年8月)。
Distribution of severity of speech disorder and condition across all speakers with more than 300 utterances recorded. For conditions, only those with > 5 speakers are shown (all others aggregated into “OTHER” for k-anonymity).
所有记录了超过300条语音的使用者中,言语障碍严重程度和状况的分布。对于状况,仅显示有>5位使用者的类别(其他类别聚合为“OTHER”以保护K-匿名性)。
ALS = amyotrophic lateral sclerosis; DS = Down syndrome; PD = Parkinson’s disease; CP = cerebral palsy; HI = hearing impaired; MD = muscular dystrophy; MS = multiple sclerosis
ALS = 肌萎缩侧索硬化症;DS = 唐氏综合征;PD = 帕金森病;CP = 脑性瘫痪;HI = 听力受损;MD = 肌营养不良症;MS = 多发性硬化症
To simplify the data collection, participants used an at-home recording system on their personal hardware (laptop or phone, with and without headphones), instead of an idealized lab-based setting that would collect studio quality recordings.
为了简化数据收集,参与者在个人设备(笔记本电脑或手机,佩戴或不佩戴耳机)上使用家庭录音系统,而非在理想化的实验室环境中采集录音室质量录音。
To reduce transcription cost, while still maintaining high transcript conformity, we prioritized scripted speech. Participants read prompts shown on a browser-based recording tool. Phrase prompts covered use-cases like home automation (“Turn on the TV.”), caregiver conversations (“I am hungry.”) and informal conversations (“How are you doing? Did you have a nice day?”). Most participants received a list of 1500 phrases, which included 1100 unique phrases along with 100 phrases that were each repeated four more times.
为降低转录成本,同时保持较高的转录一致性,我们优先采用脚本化语音。参与者阅读浏览器录音工具上显示的提示语句。短语提示涵盖多种应用场景,如家庭自动化(“打开电视”)、照护者对话(“我饿了”)和非正式对话(“你好吗?今天过得愉快吗?”)。大多数参与者收到一份包含1500个短语的列表,其中包括1100个独特短语以及100个重复四次以上的短语。
Speech professionals conducted a comprehensive auditory-perceptual speech assessment while listening to a subset of utterances for every speaker providing the following speaker-level metadata: speech disorder type (e.g., stuttering, dysarthria, apraxia), rating of 24 features of abnormal speech (e.g., hypernasality, articulatory imprecision, dysprosody), as well as recording quality assessments of both technical (e.g., signal dropouts, segmentation problems) and acoustic (e.g., environmental noise, secondary speaker crosstalk) features.
言语专业人士在听取每位使用者的部分语音时进行了全面的听觉-感知言语评估,并提供以下使用者级元数据:言语障碍类型(如口吃、构音障碍、言语失用)、24项异常言语特征的评分(如鼻音过重、发音不准确、韵律异常),以及技术特征(如信号丢失、分段问题)和声学特征(如环境噪声、次要说话者串扰)的录音质量评估。
Personalized ASR Models
个性化ASR模型
This expanded impaired speech dataset is the foundation of our new approach to personalized ASR models for disordered speech. Each personalized model uses a standard end-to-end, RNN-Transducer (RNN-T) ASR model that is fine-tuned using data from the target speaker only.
这一扩展的障碍言语数据集是我们针对障碍言语开发个性化ASR模型新方法的基础。每个个性化模型使用标准的端到端RNN-转换器(RNN-T)ASR模型,该模型仅使用目标说话者的数据进行微调。
Architecture of RNN-Transducer. In our case, the encoder network consists of 8 layers and the predictor network consists of 2 layers of uni-directional LSTM cells.
RNN-转换器的架构。在我们的案例中,编码器网络由8层组成,预测器网络由2层单向LSTM单元组成。
To accomplish this, we focus on adapting the encoder network, i.e. the part of the model dealing with the specific acoustics of a given speaker, as speech sound disorders were most common in our corpus. We found that only updating the bottom five (out of eight) encoder layers while freezing the top three encoder layers (as well as the joint layer and decoder layers) led to the best results and effectively avoided overfitting. To make these models more robust against background noise and other acoustic effects, we employ a configuration of SpecAugment specifically tuned to the prevailing characteristics of disordered speech. Further, we found that the choice of the pre-trained base model was critical. A base model trained on a large and diverse corpus of typical speech (multiple domains and acoustic conditions) proved to work best for our scenario.
为实现这一目标,我们专注于适配编码器网络,即处理特定说话者声学特征的部分,因为言语障碍在我们的语料库中最常见。我们发现,只更新底部五层(共八层)编码器,同时冻结顶部三层编码器(以及联合层和解码器层)能获得最佳结果,并有效避免过拟合。为使这些模型对背景噪声和其他声学效应更鲁棒,我们采用了针对障碍言语主要特征专门调整的SpecAugment配置。此外,我们发现预训练基础模型的选择至关重要。在大型多样化典型语音语料库(多领域和声学条件)上训练的基础模型被证明最适合我们的场景。
Results
结果
We trained personalized ASR models for ~430 speakers who recorded at least 300 utterances. 10% of utterances were held out as a test set (with no phrase overlap) on which we calculated the word error rate (WER) for the personalized model and the unadapted base model.
我们为约430位至少录制了300条语音的使用者训练了个性化ASR模型。其中10%的语音作为测试集(无短语重叠),我们在该测试集上计算了个性化模型和未适配基础模型的单词错误率(WER)。
Overall, our personalization approach yields significant improvements across all severity levels and conditions. Even for severely impaired speech, the median WER for short phrases from the home automation domain dropped from around 89% to 13%. Substantial accuracy improvements were also seen across other domains such as conversational and caregiver.
总体而言,我们的个性化方法在所有严重程度和状况下均带来了显著改进。即使对于严重障碍语音,家庭自动化领域短句的中位WER也从约89%降至13%。在其他领域(如对话和照护者对话)也观察到了显著的准确性提升。
WER of unadapted and personalized ASR models on home automation phrases.
未适配和个性化ASR模型在家庭自动化短语上的WER。
To understand when personalization does not work well, we analyzed several subgroups:
为了解个性化在哪些情况下效果不佳,我们分析了几个子组:
HighWER and LowWER: Speakers with high and low personalized model WERs based on the 1st and 5th quintiles of the WER distribution.
HighWER和LowWER:基于WER分布的第1和第5五分位数,分别具有高和低个性化模型WER的说话者。
SurpHighWER: Speakers with a surprisingly high WER (participants with typical speech or mild speech impairment of the HighWER group).
SurpHighWER:WER出乎意料高的说话者(HighWER组中具有典型语音或轻度言语障碍的参与者)。
Different pathologies and speech disorder presentations are expected to impact ASR non-uniformly. The distribution of speech disorder types within the HighWER group indicates that dysarthria due to cerebral palsy was particularly difficult to model. Not surprisingly, median severity was also higher in this group.
不同的病理和言语障碍表现预期会对ASR产生非均匀的影响。HighWER组中言语障碍类型的分布表明,脑性瘫痪引起的构音障碍尤其难以建模。毫不奇怪,该组的中位严重程度也更高。
To identify the speaker-specific and technical factors that impact ASR accuracy, we examined the differences (Cohen's D) in the metadata between the participants that had poor (HighWER) and excellent (LowWER) ASR performance. As expected, overall speech severity was significantly lower in the LowWER group than in the HighWER group (p < 0.01). Intelligibility and severity were the most prominent atypical speech features in the HighWER group; however, other speech features also emerged, including abnormal prosody, articulation, and phonation. These speech features are known to degrade overall speech intelligibility.
为识别影响ASR准确率的说话者特定因素和技术因素,我们检查了ASR表现较差(HighWER)和优异(LowWER)参与者之间元数据的差异(Cohen's D)。正如预期,LowWER组的整体言语严重程度显著低于HighWER组(p < 0.01)。可懂度和严重程度是HighWER组中最突出的非典型语音特征;然而,其他语音特征也显现出来,包括异常的韵律、发音和发声。这些语音特征已知会降低整体语音可懂度。
The SurpHighWER group had fewer training utterances and lower SNR compared with the LowWER group (p < 0.01) resulting in large (negative) effect sizes, with all other factors having small effect sizes, except fastness. In contrast, the HighWER group exhibited medium to large differences across all factors.
与LowWER组相比,SurpHighWER组的训练语音更少,信噪比(SNR)更低(p < 0.01),导致较大的(负向)效应量,除“快速”外,所有其他因素的效应量均较小。相比之下,HighWER组在所有因素上均表现出中等到较大的差异。
Speech disorder and technical metadata effect sizes for the HighWER-vs-LowWER and SurpHighWER-vs-LowWER pairs. Positive effects indicated that the group values of the HighWER group were greater than LowWER groups.
HighWER-vs-LowWER和SurpHighWER-vs-LowWER配对中言语障碍和技术元数据的效应量。正向效应表明HighWER组的组值大于LowWER组。
We then compared personalized ASR models to human listeners. Three speech professionals independently transcribed 30 utterances per speaker. We found that WERs were, on average, lower for personalized ASR models compared to the WERs of human listeners, with gains increasing by severity.
随后,我们将个性化ASR模型与人类听众进行了比较。三位言语专业人士独立转录每位说话者的30条语音。我们发现,个性化ASR模型的平均WER低于人类听众的WER,且收益随严重程度增加而增大。
Delta between the WERs of the personalized ASR models and the human listeners. Negative values indicate that personalized ASR performs better than human (expert) listeners.
个性化ASR模型与人类听众WER之间的差值。负值表示个性化ASR表现优于人类(专家)听众。
Conclusions
结论
With over 1 million utterances, Euphonia’s corpus is one of the largest and most diversely disordered speech corpora (in terms of disorder types and severities) and has enabled significant advances in ASR accuracy for these types of atypical speech. Our results demonstrate the efficacy of personalized ASR models for recognizing a wide range of speech impairments and severities, with potential for making ASR available to a wider population of users.
凭借超过100万条语音,Euphonia的语料库是规模最大、障碍类型和严重程度最多样化的障碍言语语料库之一,并为这些非典型语音的ASR准确率带来了显著进步。我们的结果证明了个性化ASR模型在识别广泛言语障碍和严重程度方面的有效性,具有将ASR提供给更广泛用户群体的潜力。
Acknowledgements
致谢
Key contributors to this project include Michael Brenner, Julie Cattiau, Richard Cave, Jordan Green, Rus Heywood, Pan-Pan Jiang, Anton Kast, Marilyn Ladewig, Bob MacDonald, Phil Nelson, Katie Seaver, Jimmy Tobin, and Katrin Tomanek. We gratefully acknowledge the support Project Euphonia received from members of many speech research teams across Google, including Françoise Beaufays, Fadi Biadsy, Dotan Emanuel, Khe Chai Sim, Pedro Moreno Mengibar, Arun Narayanan, Hasim Sak, Suzan Schwartz, Joel Shor, and many others. And most importantly, we wanted to say a huge thank you to the over 1300 participants who recorded speech samples and the many advocacy groups who helped us connect with these participants.
本项目的主要贡献者包括Michael Brenner、Julie Cattiau、Richard Cave、Jordan Green、Rus Heywood、Pan-Pan Jiang、Anton Kast、Marilyn Ladewig、Bob MacDonald、Phil Nelson、Katie Seaver、Jimmy Tobin和Katrin Tomanek。我们衷心感谢Project Euphonia从Google众多语音研究团队获得的支持,包括Françoise Beaufays、Fadi Biadsy、Dotan Emanuel、Khe Chai Sim、Pedro Moreno Mengibar、Arun Narayanan、Hasim Sak、Suzan Schwartz、Joel Shor等许多人。最重要的是,我们想向超过1300位录制语音样本的参与者以及帮助我们与这些参与者建立联系的众多倡导团体表示由衷的感谢。