Voice Command Recognition for 3D Endless Games Using Hybrid Transformer LSTM
DOI:
https://doi.org/10.34148/teknika.v15i1.1414Keywords:
Speech Recognition, BiLSTM, Transformer, Hybrid Model, MFCCAbstract
Voice command recognition has become increasingly important for enabling natural human–computer interaction in gaming and embedded systems. However, achieving accurate and noise-robust recognition on small-scale datasets remains challenging due to the limited availability of data and computational resources. To address this, this paper presents a systematic study on architectural choices for small-scale speech command recognition. We compare three neural architectures, BiLSTM, Transformer, and a sequential hybrid, in a controlled framework using identical MFCC front-ends and standardized noise augmentation. Experiments include various configurations, varying sampling rates, and noise types, rigorously evaluated using repeated cross-validation to ensure reliability. The results show that the hybrid architecture achieves superior accuracy, clearly outperforming the standalone BiLSTM and standalone Transformer baseline architectures. The hybrid model exhibits lower variance across cross-validation folds and initialization processes, and demonstrates significantly higher training throughput compared to the BiLSTM model. The model demonstrates excellent robustness to acoustic noise; variants trained with pink and white noise augmentation show comparable accuracy, confirming robust feature learning under diverse data augmentation conditions. These findings support the hypothesis that BiLSTM's local temporal modeling complements Transformer's global self-attention, enabling more effective capture of multi-scale temporal patterns in short utterances. Real-time deployment feasibility was confirmed through integration with a 3D game engine, achieving an average total response time of 5.23 ms, demonstrating the model's suitability for low-latency interactive applications.
Downloads
References
[1] D. M. Waqar, T. S. Gunawan, M. Kartiwi, and R. Ahmad, “Real-Time Voice-Controlled Game Interaction using Convolutional Neural Networks,” in 2021 IEEE 7th International Conference on Smart Instrumentation, Measurement and Applications, ICSIMA 2021, Institute of Electrical and Electronics Engineers Inc., Aug. 2021, pp. 76–81. doi: 10.1109/ICSIMA50015.2021.9526318.
[2] F. Adnan, I. Amelia, D. Sayyid ’, and U. Shiddiq, “Implementasi Voice Recognition Berbasis Machine Learning,” Edu Elektrika Journal, vol. 11, no. 1, 2022.
[3] X. Wang, P. Zhang, and C. Liu, “Acoustic signal-based identification of pipeline defects using optimized MFCC and LSTM,” Journal of Pipeline Science and Engineering, p. 100355, Sep. 2025, doi: 10.1016/j.jpse.2025.100355.
[4] K. Zaman, K. Li, M. Sah, C. Direkoglu, S. Okada, and M. Unoki, “Transformers and audio detection tasks: An overview,” Mar. 01, 2025, Elsevier Inc. doi: 10.1016/j.dsp.2024.104956.
[5] S. Ünalan, O. Günay, I. Akkurt, K. Gunoglu, and H. O. Tekin, “A comparative study on breast cancer classification with stratified shuffle split and K-fold cross validation via ensembled machine learning,” J Radiat Res Appl Sci, vol. 17, no. 4, p. 101080, Dec. 2024, doi: 10.1016/j.jrras.2024.101080.
[6] S. Riaz, A. Saghir, M. J. Khan, H. Khan, H. S. Khan, and M. J. Khan, “TransLSTM: A hybrid LSTM-Transformer model for fine-grained suggestion mining,” Natural Language Processing Journal, vol. 8, p. 100089, Sep. 2024, doi: 10.1016/j.nlp.2024.100089.
[7] A. A. Alsuwaylimi, “Arabic dialect identification in social media: A hybrid model with transformer models and BiLSTM,” Heliyon, vol. 10, no. 17, Sep. 2024, doi: 10.1016/j.heliyon.2024.e36280.
[8] A. Loubser, P. De Villiers, and A. De Freitas, “End-to-end automated speech recognition using a character based small scale transformer architecture,” Expert Syst Appl, vol. 252, Oct. 2024, doi: 10.1016/j.eswa.2024.124119.
[9] Y. Yan, S. O. Simons, L. van Bemmel, L. G. Reinders, F. M. E. Franssen, and V. Urovi, “Optimizing MFCC parameters for the automatic detection of respiratory diseases,” Applied Acoustics, vol. 228, Jan. 2025, doi: 10.1016/j.apacoust.2024.110299.
[10] A. Sabha and A. Selwal, “A novel Approach for Audio-based Video Analysis via MFCC Features,” in Procedia Computer Science, Elsevier B.V., 2024, pp. 1512–1521. doi: 10.1016/j.procs.2024.04.142.
[11] K. Pizzi, M. Pizarro, and A. Fischer, “Comparative study on noise-augmented training and its effect on adversarial robustness in ASR systems,” Comput Speech Lang, vol. 96, Feb. 2026, doi: 10.1016/j.csl.2025.101869.
[12] S. Tirronen, S. R. Kadiri, and P. Alku, “The Effect of the MFCC Frame Length in Automatic Voice Pathology Detection,” Journal of Voice, vol. 38, no. 5, pp. 975–982, Sep. 2024, doi: 10.1016/j.jvoice.2022.03.021.
[13] J. Tayebi, A. Rezaie, M. Rezaie, M. Hassanpour, and M. R. I. Faruque, “Bi-LSTM neural network for enhanced radiotherapy: Detection of tissue parameters,” Nuclear Engineering and Technology, vol. 58, no. 1, p. 103906, Jan. 2026, doi: 10.1016/j.net.2025.103906.
[14] P. Calle et al., “Integration of nested cross-validation, automated hyperparameter optimization, high-performance computing to reduce and quantify the variance of test performance estimation of deep learning models,” Comput Methods Programs Biomed, vol. 272, Dec. 2025, doi: 10.1016/j.cmpb.2025.109063.
[15] B. Perrone, F. Amato, and G. Olmo, “Voice classification in Parkinson’s disease: A deep learning approach using transformers and error rate metrics,” Biomed Signal Process Control, vol. 113, Mar. 2026, doi: 10.1016/j.bspc.2025.108954.
[16] F. Andayani, L. B. Theng, M. T. Tsun, and C. Chua, “Hybrid LSTM-Transformer Model for Emotion Recognition From Speech Audio Files,” IEEE Access, vol. 10, pp. 36018–36027, 2022, doi: 10.1109/ACCESS.2022.3163856.
[17] J. Zhao et al., “Advancing calving time prediction using tail acceleration data: Comparative evaluation of transformer, hybrid CNN-transformer and LSTM-transformer models,” Smart Agricultural Technology, vol. 12, p. 101531, Dec. 2025, doi: 10.1016/j.atech.2025.101531.
[18] H. Mou, H. Rong, and A. P. Teixeira, “Detecting abnormal ship trajectory to avoid bridge collisions via a Transformer-BiLSTM model,” Ocean Engineering, vol. 343, Jan. 2026, doi: 10.1016/j.oceaneng.2025.123232.
[19] Z. Jamshidzadeh, M. Ehteram, and H. Shabanian, “Bidirectional Long Short-Term Memory (BILSTM) - Support Vector Machine: A new machine learning model for predicting water quality parameters,” Ain Shams Engineering Journal, vol. 15, no. 3, Mar. 2024, doi: 10.1016/j.asej.2023.102510.
[20] T. Leinonen et al., “Empirical investigation of multi-source cross-validation in clinical ECG classification,” Comput Biol Med, vol. 183, Dec. 2024, doi: 10.1016/j.compbiomed.2024.109271.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Teknika

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.















