Natural and Engineering Sciences

TRANSFER LEARNING BASED UZBEK AUTOMATIC SPEECH RECOGNITION USING XLS R AND N GRAM LANGUAGE MODELING

Vol. 4 No. 8 (2026): Central Asian Journal of Academic Research 112-117

DOI: 10.5281/zenodo.21946983 2026-08-14 Natural and Engineering Sciences Open Access

Authors

  • Sharipov Zohirjon Zokirjon ugli

Abstract

Uzbek is a Turkic, morphologically rich, and comparatively low resource language for which large annotated speech datasets remain scarce. Conventional automatic speech recognition (ASR) pipelines depend on hundreds of hours of transcribed audio, which is costly to obtain for such languages. This study investigates how effectively a multilingual self supervised speech representation model  Wav2Vec2 XLS R, pretrained on roughly 436,000 hours of audio across 128 languages can be adapted to Uzbek ASR through transfer learning and shallow fusion with an n gram language model. We fine tune the pretrained XLS R encoder on Uzbek read speech data from the Mozilla Common Voice corpus using a Connectionist Temporal Classification (CTC) objective over a character vocabulary, and we integrate a KenLM n gram language model at decoding time through the Wav2Vec2ProcessorWithLM beam search decoder. System quality is measured with Word Error Rate (WER) and Character Error Rate (CER) on a held out test set, comparing three configurations: a zero shot XLS R baseline, the fine tuned acoustic model, and the fine tuned model combined with the n gram language model. Results indicate that fine tuning substantially reduces error over the baseline and that n gram fusion yields a further reduction, confirming that transfer learning is an effective strategy for low resource Uzbek ASR. The main contributions are the adaptation of XLS R to Uzbek, the integration of an n gram decoder, a WER/CER evaluation across configurations, and a deployable prototype.

Keywords:

References

Baevski, A., Zhou, H., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A Framework for Self Supervised Learning of Speech Representations. arXiv:2006.11477. https://arxiv.org/abs/2006.11477

Babu, A., et al. (2021). XLS R: Self supervised Cross lingual Speech Representation Learning at Scale. arXiv:2111.09296. https://arxiv.org/abs/2111.09296

Ardila, R., et al. (2020). Common Voice: A Massively Multilingual Speech Corpus. arXiv:1912.06670. https://arxiv.org/abs/1912.06670

Musaev, M., Mussakhojayeva, S., Khujayarov, I., Khassanov, Y., Ochilov, M., & Varol, H. A. (2021). USC: An Open Source Uzbek Speech Corpus and Initial Speech Recognition Experiments. arXiv:2107.14419. https://arxiv.org/abs/2107.14419

Graves, A., Fernandez, S., Gomez, F., & Schmidhuber, J. (2006). Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks. ICML.

Heafield, K. (2011). KenLM: Faster and Smaller Language Model Queries. Proc. WMT / ACL. https://aclanthology.org/W11 2123/

von Platen, P. (2021). Fine Tune Wav2Vec2 for English ASR with Transformers. Hugging Face Blog. https://huggingface.co/blog/fine tune wav2vec2 english

von Platen, P. (2022). Boosting Wav2Vec2 with n grams in Transformers. Hugging Face Blog. https://huggingface.co/blog/wav2vec2 with ngram

Readership

8 Views
0 PDF downloads
6 Countries

    Downloads

    Published

    2026-08-14

    Issue

    Section

    Natural and Engineering Sciences

    How to Cite

    Sharipov, Z. (2026). TRANSFER LEARNING BASED UZBEK AUTOMATIC SPEECH RECOGNITION USING XLS R AND N GRAM LANGUAGE MODELING. Central Asian Journal of Academic Research, 4(8), 112-117. https://doi.org/10.5281/zenodo.21946983
    Innovative Academy RSC
    Article metrics Views and PDF downloads
    2 Views
    0 Downloads