doi: 10.17586/2226-1494-2026-26-3-447-456


Analytical review of end-to-end speech translation methods based on acoustic-semantic representations

A. O. Ostrovskii, A. A. Karpov


Read the full article  ';
Article in Russian

For citation:
Ostrovskii A.O., Karpov A.A. Analytical review of end-to-end speech translation methods based on acoustic-semantic representations. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2026, vol. 26, no. 3, pp. 447–456 (in Russian). doi: 10.17586/2226-1494-2026-26-3-447-456


Abstract
This paper analyzes contemporary methods of end-to-end speech-to-speech translation that employ discrete representations of the speech signal as an intermediate representation. The relevance of the topic is driven by the growing demand for real-time machine speech translation systems capable of preserving the unique vocal characteristics of the speaker. The study examines approaches that perform cross-lingual speech-to-speech conversion without an intermediate transition to text. Particular attention is paid to the role of discrete units as carriers of semantic and paralinguistic information as well as their significance in end-to-end speech translation architectures. A systematic review and comparative analysis of the literature covering the period 2012–2025, encompassing more than 50 publications, was conducted. Direct speech translation architectures, self-supervised representation learning methods, neural audio codecs, and multilingual scaling systems were analyzed. For each class of methods, the design principles, advantages, and limitations were identified. The main data corpora and translation quality evaluation metrics were reviewed. It is shown that the transition from cascaded machine translation architectures to end-to-end methods reduces error accumulation and decreases translation latency. It is further demonstrated that, when pre-training is employed, the speech synthesis quality of end-to-end systems approaches that of cascaded ones. Discrete acoustic units obtained via self-supervised learning methods and neural codecs are capable of preserving the information necessary for recovering semantic content and individual speech characteristics. Key research gaps have been identified: in the majority of works, discrete representations are used as a technical tool, while their internal structure, linguistic interpretability, and information capacity remains insufficiently studied. The findings of the review may be applied in the design of simultaneous machine translation systems that preserve the vocal characteristics of the speaker. Directions for further research have been formulated, including analysis of the internal structure of discrete representations, development of methods for controlled cross-lingual transfer of paralinguistic characteristics, and the creation of metrics for separately evaluating semantic accuracy and prosody preservation quality in translation. The conclusions obtained are of interest both for fundamental research and for applied problems in the development of machine translation systems.

Keywords: speech-to-speech translation, end-to-end models, discrete speech representations, neural speech synthesis, paralinguistic features

Acknowledgements. This research was supported by the Russian state research topic of SPC RAS No. FFZF-2025-0003.

References
1. Sethiya N., Maurya C.K. End-to-end speech-to-text translation: a survey. Computer Speech & Language, 2025, vol. 90, pp. 101751. doi: 10.1016/j.csl.2024.101751
2. Bentivogli L., Cettolo M., Gaido M., Karakanta A., Martinelli A., Negri M., Turchi M. Cascade versus direct speech translation: Do the differences still make a difference? Proc. of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 2021, pp. 2873–2887. doi: 10.18653/v1/2021.acl-long.224
3. Jia Y., Weiss R.J., Biadsy F., Macherey W., Johnson M., Chen Z., Wu Y. Direct speech-to-speech translation with a sequence-to-sequence model. Proc. of the Annual Conference of the International Speech Communication Association Interspeech, 2019, pp. 1123–1127. doi: 10.21437/Interspeech.2019-1951
4. Lee A., Chen P.-J., Wang C., Gu J., Popuri S., et al. Direct speech-to-speech translation with discrete units. Proc. of the 60th Annual Meeting of the Association for Computational Linguistics, 2022, vol. 1, pp. 3327–3339. doi: 10.18653/v1/2022.acl-long.235
5. Lee A., Gong H., Duquenne P.-A., Schwenk H., Chen P.-J., Wang C., et al. Textless speech-to-speech translation on real data. Proc. of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 860–872. doi: 10.18653/v1/2022.naacl-main.63
6. Baevski A., Schneider S., Auli M. vq-wav2vec: Self-supervised learning of discrete speech representations. Proc. of the International Conference on Learning Representations, 2020.
7. Baevski A., Zhou Y., Mohamed A., Auli M. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 2020, vol. 33. P. 12449–12460.
8. Hsu W.-N., Bolte B., Tsai Y.-H.H., Lakhotia K., Salakhutdinov R., Mohamed A. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, vol. 29, pp. 3451–3460. doi: 10.1109/TASLP.2021.3122291
9. Chen S., Wang C., Chen Z., Wu Y., Liu S., Chen Z., et al. WavLM: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 2022, vol. 16, no. 6, pp. 1505–1518. doi: 10.1109/JSTSP.2022.3188113
10. Zeghidour N., Luebs A., Omran A., Skoglund J., Tagliasacchi M. SoundStream: an end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, vol. 30, pp. 495–507. doi: 10.1109/TASLP.2021.3129994
11. Défossez A., Copet J., Synnaeve G., Adi Y. High fidelity neural audio compression. arXiv, 2022. arXiv:2210.13438. doi: 10.48550/arXiv.2210.13438
12. Chen S., Wang C., Wu Y., Zhang Z., Zhou L., Liu S., et al. Neural codec language models are zero-shot text to speech synthesizers. IEEE Transactions on Audio, Speech and Language Processing, 2025, vol. 33, pp. 705–718. doi: 10.1109/TASLPRO.2025.3530270
13. Inaguma H., Popuri S., Kulikov I., Chen P.-J., Wang C., Chung Y.-A., et al. UnitY: Two-pass direct speech-to-speech translation with discrete units. Proc. of the 61st Annual Meeting of the Association for Computational Linguistics, 2023, pp. 15655–15680. doi: 10.18653/v1/2023.acl-long.872
14. SeamlessM4T: Massively multilingual & multimodal machine translation. arXiv, 2023. arXiv:2308.11596. doi: 10.48550/arXiv.2308.11596
15. Jia Y., Ramanovich M.T., Remez T., Pomerantz R. Translatotron 2: High-quality direct speech-to-speech translation with voice preservation. Proc. of the 39th International Conference on Machine Learning, 2022, vol. 162, pp. 10120–10134.
16. Nachmani E., Levkovitch A., Ding Y., Asawaroengchai C., Zen H., Ramanovich M.T. Translatotron 3: Speech to speech translation with monolingual data. Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10686–10690. doi: 10.1109/ICASSP48485.2024.10448426
17. Costa-jussà M.R., Cross J., Çelebi O., Elbayad M., Heafield K., Heffernan K. et al., No language left behind: Scaling human-centered machine translation. arXiv, 2022. arXiv:2207.04672. doi: 10.48550/arXiv.2207.04672
18. Barrault L., Chung Y.-A., Meglioli M.C., Dale D., Dong N., Duppenthaler M., et al. Seamless: Multilingual expressive and streaming speech translation. arXiv, 2023. arXiv:2312.05187. doi: 10.48550/arXiv.2312.05187
19. Karpov A.A., Verkhodanova V.O. Speech technologies for under-resourced languages of the world. Voprosy Jazykoznanija, 2015, no. 2, pp. 117–135.(in Russian)
20. Rubenstein P.K., Asawaroengchai C., Nguyen D.D., Bapna A., Borsos Z., de Chaumont Quitry F., et al. AudioPaLM: A large language model that can speak and listen. arXiv, 2023. arXiv:2306.12925. doi: 10.48550/arXiv.2306.12925
21. Borsos Z., Marinier R., Vincent D., Kharitonov E., Pietquin O., Sharifi M., et al. AudioLM: A language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023, vol. 31, pp. 2523–2533. doi: 10.1109/TASLP.2023.3288409
22. Zhang D., Li S., Zhang X., Zhan J., Wang P., Zhou Y., et al. SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities. Findings of the Association for Computational Linguistics: EMNLP, 2023, pp. 15757–15773. doi: 10.18653/v1/2023.findings-emnlp.1055
23. Zeng A., Du Z., Liu M., Wang K., Jiang S., Zhao L., et al. GLM-4-Voice: Towards intelligent and human-like end-to-end spoken chatbot. arXiv, 2024. arXiv:2412.02612. 2024. doi: 10.48550/arXiv.2412.02612
24. Kahn J., Rivière M., Zheng W., Kharitonov E., Xu Q., Mazaré P.-E., et al. Libri-light: A benchmark for ASR with limited or no supervision. Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 7669–7673. doi: 10.1109/ICASSP40776.2020.9052942
25. Ardila R., Branson M., Davis K., Kohler M., Meyer J., Henretty M., et al. Common Voice: A massively-multilingual speech corpus. Proc. of the 12th Language Resources and Evaluation Conference (LREC), 2020, pp. 4218–4222.
26. Pratap V., Xu Q., Sriram A., Synnaeve G., Collobert R. MLS: A large-scale multilingual dataset for speech research. Proc. of the Annual Conference of the International Speech Communication Association Interspeech, 2020, pp. 2757–2761. doi: 10.21437/Interspeech.2020-2826
27. Wang C., Rivière M., Lee A., Wu A., Talnikar C., Haziza D., et al. VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. Proc. of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 2021, pp. 993–1003. doi: 10.18653/v1/2021.acl-long.80
28. Wang C., Wu A., Gu J., Pino J. CoVoST 2 and massively multilingual speech translation. Proc. of the Annual Conference of the International Speech Communication Association Interspeech, 2021, pp. 2247–2251. doi: 10.21437/Interspeech.2021-2027
29. Jia Y., Ramanovich M.T., Wang Q., Zen H. CVSS corpus and massively multilingual speech-to-speech translation. Proc. of the 13th Language Resources and Evaluation Conference (LREC), 2022, pp. 6691–6703.
30. Di Gangi M.A., Cattoni R., Bentivogli L., Negri M., Turchi M. MuST-C: a multilingual speech translation corpus. Proc. of the 19th Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2019, pp. 2012–2017. doi: 10.18653/v1/N19-1202
31. Conneau A., Ma M., Khanuja S., Zhang Y., Axelrod V., Dalmia S., et al. FLEURS: Few-shot learning evaluation of universal representations of speech. Proc. of the IEEE Spoken Language Technology Workshop (SLT), 2022, pp. 798–805. doi: 10.1109/SLT54892.2023.10023141
32. Chen M., Duquenne P.-A., Andrews P., Kao J., Mourachko A., Schwenk H., et al. BLASER: A text-free speech-to-speech translation evaluation metric. Proc. of the 61st Annual Meeting of the Association for Computational Linguistics, 2023, pp. 9064–9079. doi: 10.18653/v1/2023.acl-long.504
33. Saeki T., Xin D., Nakata W., Koriyama T., Takamichi S., Saruwatari H. UTMOS: UTokyo-SaruLab system for VoiceMOS Challenge 2022. Proc. of the Annual Conference of the International Speech Communication Association Interspeech, 2022, pp. 4521–4525. doi: 10.21437/Interspeech.2022-439
34. Kim J., Kong J., Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. Proc. of the International Conference on Machine Learning, 2021,vol. 139, pp. 5530–5540.
35. Ivanko D.V., Ryumin D.A. Automatic sign language translation: a review of neural network methods for recognition and synthesis of spoken and signed language. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2024, vol. 24, no. 5, pp. 669–686. (in Russian). doi: 10.17586/2226-1494-2024-24-5-669-686


Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License
Copyright 2001-2026 ©
Scientific and Technical Journal
of Information Technologies, Mechanics and Optics.

Яндекс.Метрика