doi: 10.17586/2226-1494-2026-26-4-673-682


Deep learning for speech enhancement: architectures, paradigms, and emerging trends

D. V. Ivanko


Read the full article  ';
Article in Russian

For citation:
Ivanko D.V. Deep learning for speech enhancement: architectures, paradigms, and emerging trends. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2026, vol. 26, no. 4, pp. 673–682 (in Russian). doi: 10.17586/2226-1494-2026-26-4-673-682


Abstract
A systematic review of modern neural network methods for Speech Enhancement is presented, aimed at improving speech intelligibility and quality under acoustic distortions. The reviewed methods can be applied in voice control systems, telecommunications, hearing aids, and human–machine interaction interfaces. Key architectural approaches are considered, including classical recurrent and convolutional networks as well as modern hybrid architectures with attention mechanisms (Transformer, Conformer), state-space models (Mamba), and advanced recurrent blocks (xLSTM). The advantages and disadvantages of different architectures are shown in terms of speech restoration quality and computational efficiency. Specific problems of existing methods are highlighted, including high computational cost and insufficient generalization capability under non-stationary noise conditions. The need for further research in the development of lightweight models for mobile devices, multi-distortion suppression methods, and the integration of neural network noise suppression with generative models to achieve a new level of speech signal restoration quality is demonstrated.

Keywords: speech enhancement, noise suppression, deep learning, neural network architectures, transformers, Mamba, xLSTM, survey

Acknowledgements. The work in the “New Paradigm for Assessing SE Systems” section was completed under budget topic No. FFZF-2025-0003. All other sections were supported by RSF grant No. 25-71-00093.

References
1. Wang H., Wang D.L. Cross-domain diffusion based speech enhancement for very noisy speechюProc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5. doi: 10.1109/ICASSP49357.2023.10096985
2. Richter J., Welker S., Lemercier J.M., Lay B., Gerkmann T. Speech enhancement and dereverberation with diffusion-based generative models. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023, vol. 31, pp. 2351–2364. doi: 10.1109/TASLP.2023.3285241
3. Zaburdaev A., Ivanko D., Ryumin D. CrossMP-SENet: transformer-based cross-attention for joint magnitude-phase speech enhancement. Lecture Notes in Computer Science, 2026, vol. 16188, pp. 174–188. doi: 10.1007/978-3-032-07959-6_13
4. Upadhyay N., Karmakar A. Speech enhancement using spectral subtraction-type algorithms: A comparison and simulation study. Procedia Computer Science, 2015, vol. 54, pp. 574–584. doi: 10.1016/j.procs.2015.06.066
5. Abd El-Fattah M.A., Dessouky M.I., Abbas A.M.,Diab S.M., El-Rabaie E.M., Al-Nuaimy W.,et al. Speech enhancement with an adaptive Wiener filter. International Journal of Speech Technology, 2014, vol. 17, no, 1, pp. 53–64. doi: 10.1007/s10772-013-9205-5
6. Ephraim Y., van Trees H.L. A signal subspace approach for speech enhancement. IEEE Transactions on Speech and Audio Processing, 1995, vol. 3, no. 4, pp. 251–266. doi: 10.1109/89.397090
7. Horev A.A., Dvoryankin S.V., Kozlachkov S.B., Vasilevskaya N.V. The analysis of the potential capabilities of methods of noise reduction and reconstruction of acoustic speech signals masked by various types of noise. Voprosy Kiberbezopasnosti, 2024, no. 1 (59), pp. 89–100. (in Russian). doi: 10.21681/2311-3456-2024-1-89-100
8. Pascual S., Bonafonte A., Serrà J. SEGAN: Speech Enhancement Generative Adversarial Network. Interspeech, 2017, pp. 3642–3646. doi: 10.21437/Interspeech.2017-1428
9. Fu S.-W., Tsao Y., Lu X., Kawai H. Raw waveform-based speech enhancement by fully convolutional networks. Proc. of the Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2017, pp. 6–12. doi: 10.1109/APSIPA.2017.8281993
10. Axyonov A., Ryumin D., Ivanko D., Kashevnik A., Karpov A.Audio-visual speech recognition in-the-wild: Multi-angle vehicle cabin corpus and attention-based method. Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 8195–8199. doi: 10.1109/ICASSP48485.2024.10448048
11. Ivanko D., Karpov A., Ryumin D., Kipyatkova I.,Saveliev A., Budkov V.,et al. Using a high-speed video camera for robust audio-visual speech recognition in acoustically noisy conditions. Lecture Notes in Computer Science, 2017, vol. 10458, pp. 757–766. doi: 10.1007/978-3-319-66429-3_76
12. Pandey A., Wang D.L. A new framework for CNN-based speech enhancement in the time domain. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2019, vol. 27, no. 7, pp. 1179–1188. doi: 10.1109/TASLP.2019.2913512
13. Ivanko D., Ryumin D., Axyonov A., Kashevnik A. Speaker-dependent visual command recognition in vehicle cabin: methodology and evaluation. Lecture Notes in Computer Science, 2021, pp. 291–302. doi: 10.1007/978-3-030-87802-3_27
14. Pandey A., Wang D.L. Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain. Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 6629–6633. doi: 10.1109/ICASSP40776.2020.9054536
15. Stoller D., Ewert S., Dixon S. Wave-u-net: A multi-scale neural network for end-to-end audio source separation. arXiv, 2018. arXiv:1806.03185. doi: 10.48550/arXiv.1806.03185
16. Lan C., Jiang J., Zhang L., Zeng Z. Blind source separation based on improved Wave-U-Net network. IEEE Access, 2023, vol. 11, pp. 125951–125958. doi: 10.1109/ACCESS.2023.3330160
17. Luo Y., Mesgarani N. Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2019, vol. 27, no. 8, pp. 1256–1266. doi: 10.1109/TASLP.2019.2915167
18. Lee D., Kim S., Choi J.-W. Inter-channel Conv-TasNet for multichannel speech enhancement. arXiv, 2021. arXiv:2111.04312. doi: 10.48550/arXiv.2111.04312
19. Zheng C., Zhang H., Liu W., Luo X., Li A., Li X., et al. Sixty years of frequency-domain monaural speech enhancement: From traditional to deep learning methods. Trends in Hearing, 2023, vol. 27, pp. 1–52. doi: 10.1177/23312165231209913
20. Wang J., Saleem N., Gunawan T.S. Towards efficient recurrent architectures: A deep LSTM neural network applied to speech enhancement and recognition. Cognitive Computation, 2024, vol. 16, no. 3, pp. 1221–1236. doi: 10.1007/s12559-024-10288-y
21. Pandey A., Wang D.L. Self-attending RNN for speech enhancement to improve cross-corpus generalization. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022, vol. 30, pp. 1374–1385. doi: 10.1109/taslp.2022.3161143
22. Ivanko D., Ryumin D., Kashevnik A., et al. DAVIS: Driver's Audio-Visual Speech recognition. Proc. of the Annual Conference of the International Speech Communication Association Interspeech, 2022, pp. 1141–1142.
23. Saleem N., Gao J., Khattak M.I., Rauf H.T., Kadry S., Shafi M. DeepResGRU: Residual gated recurrent neural network-augmented Kalman filtering for speech enhancement and recognition. Knowledge-Based Systems, 2022, vol. 238, pp. 107914. doi: 10.1016/j.knosys.2021.107914
24. Axyonov A.A., Ryumina E.V., Ryumin D.A., Ivanko D.V., Karpov A.A. Neural network-based method for visual recognition of driver’s voice commands using attention mechanism. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2023, vol. 23, no. 4, pp. 767–775. (in Russian). doi: 10.17586/2226-1494-2023-23-4-767-775
25. Beck M., Pöppel K., Spanring M., Auer A., Prudnikova O., Kopp M., et al. xLSTM: extended long short-term memory. Proc. of the 38th International Conference on Neural Information Processing Systems, 2024, pp. 107547–107603.
26. Zhang Q., Chen M., Song Z., Liu H., Zhang X., et al. Long-context modeling networks for monaural speech enhancement: a comparative study. Proc. of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2025, pp. 1–5. doi: 10.1109/WASPAA66052.2025.11230983
27. Kühne N.L., Østergaard J., Jensen J., Tan Z.-H. xLSTM-SENet: xLSTM for single-channel speech enhancemen. Proc. of the Annual Conference of the International Speech Communication Association Interspeech, 2025, doi: 10.21437/interspeech.2025-108
28. Saleem N., Gunawan T.S., Kartiwi M., Nugroho B.S., Wijayanto I.NSE-CATNet: Deep neural speech enhancement using convolutional attention transformer network. IEEE Access, 2023, vol. 11, pp. 66979–66994. doi: 10.1109/ACCESS.2023.3290908
29. Yu W., Zhou J., Wang H.B., Tao L.SETransformer: Speech enhancement transformer // Cognitive Computation, 2022, vol. 14, no. 3, pp. 1152–1158. doi: 10.1007/s12559-020-09817-2
30. Jannu C., Vanambathina S.D. Convolutional transformer based local and global feature learning for speech enhancement. International Journal of Advanced Computer Science and Applications, 2023, vol. 14, no. 1, pp. 81.doi: 10.14569/IJACSA.2023.0140181
31. Lependin A.A., Nasretdinov R.S., Ilyashenko I.D. Speech enhancement method based on modified encoder-decoder pyramid transformer. Proceedings of the Institute for System Programming of the RAS, 2022, vol. 34, no. 4, pp. 135–152. (in Russian). doi: 10.15514/ISPRAS-2022-34(4)-10
32. Li M., Liu Y., Zhou L. DeConformer-SENet: An efficient deformable conformer speech enhancement network. Digital Signal Processing, 2025, vol. 156, part A, pp. 104787. doi: 10.1016/j.dsp.2024.104787
33. Xu X., Tu W., Yang Y. Pcnn: A lightweight parallel conformer neural network for efficient monaural speech enhancement. Proc. of the Annual Conference of the International Speech Communication Association Interspeech, 2023, doi: 10.21437/interspeech.2023-1376
34. Rekesh D., Koluguri N.R., Kriman S., Majumdar S., Noroozi V., Huang H., et al. Fast conformer with linearly scalable attention for efficient speech recognition. Proc. of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2023., pp. 1–8. doi: 10.1109/asru57964.2023.10389701
35. Gu A., Dao T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv, 2023. arXiv:2312.00752. doi: 10.48550/arXiv.2312.00752
36. Chao R., Cheng W.-H., La Quatra M., Siniscalchi S.M., Huck Yang C.-H., Fu S.-W., et al. An investigation of incorporating mamba for speech enhancement. Proc. of the IEEE Spoken Language Technology Workshop (SLT), 2024, pp. 302–308. doi: 10.1109/slt61566.2024.10832332
37. Wang J., Lin Z., Wang T., Ge M., Wang L., Dang J. Mamba-SEUNet: Mamba UNet for monaural speech enhancement. Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1–5. doi: 10.1109/ICASSP49660.2025.10889525
38. Chen M., Zhang Q., Wang M., Zhang X., Liu H., Ambikairaiah E., et al. Selective state space model for monaural speech enhancement. IEEE Transactions on Consumer Electronics, 2025, vol. 71, no. 2, pp. 5414–5424. doi: 10.1109/TCE.2024.3523297
39. Zhang X., Zhang Q., Liu H., Xiao T., Qian X., Ahmed B., et al. Mamba in speech: Towards an alternative to self-attention. IEEE Transactions on Audio, Speech and Language Processing, 2025, vol. 33, pp. 1933–1948. doi: 10.1109/TASLPRO.2025.3566210
40. Chung H., Plourde E., Champagne B. Discriminative training of NMF model based on class probabilities for speech enhancement. IEEE Signal Processing Letters, 2016, vol. 23, no. 4, pp. 502–506. doi: 10.1109/LSP.2016.2532903
41. Li C., Cornell S., Watanabe S., Qian Y. Diffusion-based generative modeling with discriminative guidance for streamable speech enhancement. Proc. of the IEEE Spoken Language Technology Workshop (SLT), 2024, pp. 333–340. doi: 10.1109/SLT61566.2024.10832147
42. Donahue C., Li B., Prabhavalkar R. Exploring speech enhancement with generative adversarial networks for robust speech recognition. Proc. of the IEEE international conference on acoustics, speech and signal processing (ICASSP), 2018, pp. 5024–5028.  doi: 10.1109/ICASSP.2018.8462581
43. Lu Y.J., Wang Z.-Q., Watanabe S., Richard A., Yu C., Tsao Y. Conditional diffusion probabilistic model for speech enhancement. Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 7402–7406. doi: 10.1109/ICASSP43922.2022.9746901
44. Yen H., Germain F.G., Wichern G., Le Roux J. Cold diffusion for speech enhancement. Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5. doi: 10.1109/ICASSP49357.2023.10096064
45. Serrà J., Pascual S., Pons J., Oguz Araz R., Scaini D. Universal speech enhancement with score-based diffusion. arXiv, 2022. arXiv:2206.03065. doi: 10.48550/arXiv.2206.03065
46. Lu Y.-J., Tsao Y., Watanabe S. A study on speech enhancement based on diffusion probabilistic model. Proc. of the Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2021, pp. 659–666.
47. Saijo K., Zhang W., Cornell S., Scheibler R., Li C., Ni Z., et al. Interspeech 2025 URGENT speech enhancement challenge. Proc. of the Annual Conference of the International Speech Communication Association Interspeech, 2025, doi: 10.21437/interspeech.2025-1363
48. Dubey H., Aazami A., Gopal V., Naderi B., Braun S., Cutler R., et al. ICASSP 2023 deep noise suppression challenge. IEEE Open Journal of Signal Processing, 2024, vol. 5, pp. 725–737. doi: 10.1109/OJSP.2024.3378602
49. Watanabe S., Mandel M., Barker J., Vincent E., Arora A., Chang X., et al. CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings. arXiv, 2020. arXiv:2004.09249. doi: 10.48550/arXiv.2004.09249


Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License
Copyright 2001-2026 ©
Scientific and Technical Journal
of Information Technologies, Mechanics and Optics.

Яндекс.Метрика