A Novel Deep Hybrid Parallel Model of CNN- BiLSTM-Transformer for Speech Emotion Recognition
DOI:
https://doi.org/10.32985/ijeces.17.8.1Keywords:
Emotion Recognition, Voice, Deep Learning, Signal Processing, CNN, Bidirectional LSTM, TransformerAbstract
Human speech carries not only words but also rich paralinguistic signals that reflect emotional states. Despite significant progress in automatic speech emotion recognition (SER), many existing models still struggle to fully exploit vocal information. In this paper, we propose a novel deep hybrid model with three parallel structures: Convolutional Neural Networks (CNNs), Bi-directional Long Short-Term Memory Networks (BiLSTMs), and Transformers. Our model extracts local, temporal, and contextual information separately. This parallel architecture retains completeness and mitigates potential information loss in sequential architectures. Tested on five datasets for classifying discrete emotion types (neutral, happy, sad, angry, fearful, disgust, surprised, and calm), our model performs outstandingly, with a classification accuracy of 99.89% for RAVDESS, 100% for TESS, 99.82% for SAVEE, 95.64% for CREMA-D, and 100% for the EMO-DB dataset, consistently outperforming recent advanced methods.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 International Journal of Electrical and Computer Engineering Systems

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.