Menu
Publications
2026
2025
2024
2023
2022
2021
2020
2019
2018
2017
2016
2015
2014
2013
2012
2011
2010
2009
2008
2007
2006
2005
2004
2003
2002
2001
Editor-in-Chief
Nikiforov
Vladimir O.
D.Sc., Prof.
Partners
doi: 10.17586/2226-1494-2026-26-4-826-834
Karelian speech recognition system with support for Karelian-Russian code-switching
Read the full article
Article in Russian
For citation:
Abstract
For citation:
Kipyatkova I.S., Dolgushin M.D., Kiseleva K.O., Kagirov I.A. Karelian speech recognition system with support for Karelian-Russian code-switching. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2026, vol. 26, no. 4, pp. 826–834 (in Russian). doi: 10.17586/2226-1494-2026-26-4-826-834
Abstract
This paper focuses on the development of an automatic speech recognition system for the Livvi-Karelian variety of the Karelian language, as it is spoken under conditions of code-switching between Karelian and Russian. The study of bilingual speech recognition methods is carried out. In order to improve the quality of speech recognition, a methodology for training text data augmentation via partial translation and intra-word code-switched wordforms artificial synthesis was developed. Acoustic modeling was performed by fine-tuning a pre-trained multilingual Wav2Vec2-BERT 2.0 model with the use of the data from two previously collected corpora containing 7.5 hours of speech. Fine-tuning was performed using the Transformers framework. When developing the language model, in order to address the problem of limited code-switching data, an augmentation method was applied based on partial automatic translation of Karelian texts into Russian, followed by the generation of word-forms with intra-word code-switching based on special linguistic rules. On the base of formulated rules, a list of words with intra-word code-switching was generated for a language model. The experiments showed that using a full vocabulary that includes generated hybrid word forms yields a consistent improvement in results. A further reduction in word error rate to 25.82 % on the development set and 29 % on the test portion of the corpus was achieved through linear interpolation of the Karelian language model with the Russian language model (interpolation weight 0.7). The conducted experiments confirm the effectiveness of the developed methodology for developing a bilingual speech recognition system. In particular, it is recommended to combine fine-tuning of multilingual acoustic models, text augmentation with morphological rules, and language model interpolation. The proposed approach can be applied to developing speech recognition systems for other low-resource languages of Russia spoken in an unbalanced bilingual environment.
Keywords: automatic speech recognition, code-switching, Karelian language, low-resource languages, data augmentation, intra-word code-switching, Wav2Vec2-BERT 2.0, language modeling
Acknowledgements. This research is financially supported by the Russian Science Foundation, project No. 24-21-00276, https://rscf.ru/ project/24-21-00276/.
Acknowledgements. This research is financially supported by the Russian Science Foundation, project No. 24-21-00276, https://rscf.ru/ project/24-21-00276/.

