A Nepali-Accented English Evaluation Dataset for Automatic Speech Recognition

Authors

  • Santosh Dahal Department of Electronics and Computer Engineering, Institute of Engineering, Thapathali Campus, Tribhuvan University, Nepal
  • Kiran Chandra Dahal Department of Electronics and Computer Engineering, Institute of Engineering, Thapathali Campus, Tribhuvan University, Nepal

Keywords:

automatic speech recognition, Nepali-accented English, accented speech, word error rate

Abstract

Automatic speech recognition (ASR) systems perform strongly on native-English benchmarks, yet their accuracy degrades sharply when the input speech comes from under represented non-native accents. Nepali-accented English is particularly under-served: existing resources either focus on native Nepali speech, cover broader multi-accent settings without dedicated Nepali evaluation, or provide only limited Nepali-accent coverage. This paper presents a Nepali-accented English evaluation dataset designed to support robust ASR benchmarking under accent mismatch. The corpus was collected through a web-based platform that did not collect directly identifying metadata and contains recordings from 57 speakers. Each session follows a fixed 22-prompt protocol consisting of 11 phonetic prompts, 10 domain prompts, and 1 spontaneous prompt, providing complementary coverage of pronunciation, topical vocabulary, and natural speaking style. In addition to transcribed speech, the dataset includes participant metadata for coarse exploratory subgroup analysis and speaker-level manual recording-quality labels. Manual quality assessment shows that 50.9% of sessions are clean and 42.1% contain only mild noise. As a descriptive reference, open-source ASR baselines are substantially worse on this corpus than the corresponding LibriSpeech test-clean values reported in official model cards, reaching 38.15–55.00% WER on the collected set versus reported 2–4% WER on LibriSpeech test-clean. These baseline results position the corpus as a practical held-out resource for evaluating accent robustness and out-of-distribution generalization on Nepali-accented English.

Abstract
22
PDF
18

Downloads

Published

2026-07-31

How to Cite

Dahal, S., & Dahal, K. C. (2026). A Nepali-Accented English Evaluation Dataset for Automatic Speech Recognition. Everest Advances in Science and Technology, 2(1), 55-62. https://doi.org/10.3126/east.v2i1.98619

Issue

Section

Articles

How to Cite

Dahal, S., & Dahal, K. C. (2026). A Nepali-Accented English Evaluation Dataset for Automatic Speech Recognition. Everest Advances in Science and Technology, 2(1), 55-62. https://doi.org/10.3126/east.v2i1.98619