The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic Speech

Ali, Ahmed; Shon, Suwon; Samih, Younes; Mubarak, Hamdy; Abdelali, Ahmed; Glass, James; Renals, Steve; Choukri, Khalid

Repository landing page

oai:pure.ed.ac.uk:publications/c6867a0f-50ec-42a4-958a-7e117bb40b15

The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic Speech

Authors: Ahmed Ali
Suwon Shon
Younes Samih
Hamdy Mubarak
Ahmed Abdelali
James Glass
Steve Renals
Khalid Choukri
Publication date: 20 February 2020
Publisher: Institute of Electrical and Electronics Engineers (IEEE)
Doi

Abstract

This paper describes the fifth edition of the Multi-Genre Broadcast Challenge (MGB-5), an evaluation focused on Arabic speech recognition and dialect identification. MGB-5 extends the previous MGB-3 challenge in two ways: first it focuses on Moroccan Arabic speech recognition; second the granularity of the Arabic dialect identification task is increased from 5 dialect classes to 17, by collecting data from 17 Arabic speaking countries. Both tasks use YouTube recordings to provide a multi-genre multi-dialectal challenge in the wild. Moroccan speech transcription used about 13 hours of transcribed speech data, split across training, development, and test sets, covering 7-genres: comedy, cooking, family/kids, fashion, drama, sports, and science (TEDx). The fine-grained Arabic dialect identification data was collected from known YouTube channels from 17 Arabic countries. 3,000 hours of this data was released for training, and 57 hours for development and testing. The dialect identification data was divided into three sub-categories based on the segment duration: short (under 5s), medium (5–20s), and long (>20s). Overall, 25 teams registered for the challenge, and 9 teams submitted systems for the two tasks. We outline the approaches adopted in each system and summarize the evaluation results

Similar works

Full text

Open in the Core reader

Download PDF

Edinburgh Research Explorer

oai:pure.ed.ac.uk:publications...

Last time updated on 11/05/2020

This paper was published in Edinburgh Research Explorer.

Having an issue?

Is data on this page outdated, violates copyrights or anything else? Report the problem now and we will take corresponding actions after reviewing your request.