Low-resource automatic speech recognition and error analyses of oral cancer speech

Halpern, B.M.; Feng, S.; van Son, R.; van den Brekel, M.; Scharenborg, O.

Repository landing page

oai:dare.uva.nl:openaire_cris_publications/3ef4097c-670f-4cb3-a393-b70f40d19f8b

Low-resource automatic speech recognition and error analyses of oral cancer speech

Authors: B.M. Halpern
S. Feng
R. van Son
M. van den Brekel
O. Scharenborg
Publication date: 1 June 2022
Publisher
Doi

Abstract

In this paper, we introduce a new corpus of oral cancer speech and present our study on the automatic recognition and analysis of oral cancer speech. A two-hour English oral cancer speech dataset is collected from YouTube. Formulated as a low-resource oral cancer ASR task, we investigate three acoustic modelling approaches that previously have worked well with low-resource scenarios using two different architectures; a hybrid architecture and a transformer-based end-to-end (E2E) model: (1) a retraining approach; (2) a speaker adaptation approach; and (3) a disentangled representation learning approach (only using the hybrid architecture). The approaches achieve a (1) 4.7% (hybrid) and 7.5% (E2E); (2) 7.7%; and (3) 2.0% absolute word error rate reduction, respectively, compared to a baseline system which is not trained on oral cancer speech. A detailed analysis of the speech recognition results shows that (1) plosives and certain vowels are the most difficult sounds to recognise in oral cancer speech — this problem is successfully alleviated by our proposed approaches; (3) however these sounds are also relatively poorly recognised in the case of healthy speech with the exception of/p/. (2) recognition performance of certain phonemes is strongly data-dependent; (4) In terms of the manner of articulation, E2E performs better with the exception of vowels — however, vowels have a large contribution to overall performance. As for the place of articulation, vowels, labiodentals, dentals and glottals are better captured by hybrid models, E2E is better on bilabial, alveolar, postalveolar, palatal and velar information. (5) Finally, our analysis provides some guidelines for selecting words that can be used as voice commands for ASR systems for oral cancer speakers.</p

article

Similar works

Full text

Open in the Core reader

Download PDF

International Migration, Integration and Social Cohesion online publications

oai:dare.uva.nl:openaire_cris_...

Last time updated on 08/03/2023

This paper was published in International Migration, Integration and Social Cohesion online publications.

Having an issue?

Is data on this page outdated, violates copyrights or anything else? Report the problem now and we will take corresponding actions after reviewing your request.