CNN-LSTM with Multi-Acoustic Features for Automatic Tajweed Mad Rule Classification
Abstract
The rules of mad recitation in the Qur’an are a crucial aspect of tajwīd, governing the lengthening of vowel sounds that affect both meaning and recitational accuracy. Despite its importance, there is currently no reliable automatic system capable of classifying mad rules based on voice input. This study proposes a deep learning-based approach using a hybrid Convolutional Neural Network–Long Short-Term Memory (CNN-LSTM) model to automatically classify mad rules from Qur’anic recitations. The research follows the CRISP-DM methodology, covering data understanding, preparation, modeling, and evaluation stages. Acoustic features were extracted from 3,816 annotated audio segments of Surah Al-Fātiḥah, combining Mel-Frequency Cepstral Coefficients (MFCC), Chroma, Spectral Contrast, and Root Mean Square (RMS) to represent phonetic and prosodic attributes. The CNN layers captured spatial characteristics of the spectrum, while LSTM layers modeled temporal dependencies of the audio. Experimental results show that the combination of all four features achieved an accuracy of 97.21%, precision of 95.28%, recall of 95.22%, and F1-score of 95.25%. These findings indicate that multi-feature integration enhances model robustness and interpretability. The proposed CNN-LSTM framework demonstrates potential for practical deployment in voice-based tajwīd learning tools and contributes to the broader field of Qur’anic speech recognition by offering a systematic, ethically grounded, and data-driven approach to mad classification.
Keywords
Full Text:
PDFReferences
N. Anggraini, A. Kurniawan, L. K. Wardhani, and N. Hakiem, “Speech recognition application for the speech impaired using the android-based google cloud speech API,” TELKOMNIKA (Telecommunication Computing Electronics and Control), vol. 16, no. 6, pp. 2733–2739, 2018.
Z. Liu, Y. Wang, and T. Chen, “Audio Feature Extraction and Analysis for Scene Segmentation and Classification,” The Journal of VLSI Signal Processing, vol. 20, no. 1/2, pp. 61–79, 1998, doi: 10.1023/A:1008066223044.
M. Turab, T. Kumar, M. Bendechache, and T. Saber, “Investigating multi-feature selection and ensembling for audio classification,” arXiv preprint arXiv:2206.07511, 2022.
R. Li, B. Yin, Y. Cui, Z. Du, and K. Li, “Research on Environmental Sound Classification Algorithm Based on Multi-feature Fusion,” in 2020 IEEE 9th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), IEEE, Dec. 2020, pp. 522–526. doi: 10.1109/ITAIC49862.2020.9338926.
M. A. Hossan, S. Memon, and M. A. Gregory, “A novel approach for MFCC feature extraction,” in 2010 4th international conference on signal processing and communication systems, IEEE, 2010, pp. 1–5.
S. Ewert, “Chroma Toolbox: MATLAB implementations for extracting variants of chroma-based audio features,” in Proc. ISMIR, 2011.
M. L. Massar, M. Fickus, E. Bryan, D. T. Petkie, and A. J. Terzuoli, “Fast computation of spectral centroids,” Adv Comput Math, vol. 35, no. 1, pp. 83–97, Jul. 2011, doi: 10.1007/s10444-010-9167-y.
R. Nekoutabar, F. S. Ghaheri, and H. Jalilvand, “The Effect of Root-Mean-Square and Loudness-Based Calibration Approach on the Acceptable Noise Level,” Auditory and Vestibular Research, Oct. 2024, doi: 10.18502/avr.v33i4.16654.
O. Theobald, Machine learning for absolute beginners: a plain English introduction, vol. 157. Scatterplot press London, 2017.
G. Van Houdt, C. Mosquera, and G. Nápoles, “A review on the long short-term memory model,” Artif Intell Rev, vol. 53, no. 8, pp. 5929–5955, Dec. 2020, doi: 10.1007/s10462-020-09838-1.
Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou, “A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects,” IEEE Trans Neural Netw Learn Syst, vol. 33, no. 12, pp. 6999–7019, Dec. 2022, doi: 10.1109/TNNLS.2021.3084827.
M. Maryamah, N. J. K. Pradiptamurty, H. K. Shafro, M. S. Al Qurtubi, and G. A. Tambahjong, “Speech Emotion Recognition (SER) dengan Metode Bidirectional LSTM,” in PROSIDING SEMINAR NASIONAL SAINS DATA, 2023, pp. 153–161.
R. Hadiyansah and R. Andamira, “Convolutional Neural Network (CNN) for Detecting Al-Qur’an Reciting and Memorizing,” Khazanah Journal of Religion and Technology, vol. 1, no. 2, pp. 44–48, 2023.
N. Anggraini, Zulkifli, Y. Rahman, A. N. Hidayanto, and H. T. Sukmana, “Modeling Madd Reading Classification in Surah Al-Fatihah with MFCC Feature Extraction and LSTM Algorithm,” in 2024 12th International Conference on Cyber and IT Service Management (CITSM), IEEE, Oct. 2024, pp. 1–7. doi: 10.1109/CITSM64103.2024.10775370.
V Anupama, Ch Amrutha, G Amrutha Varshini, G Sai Gowtham Nandan, and GVLN Satya Sai Vivek, “A MFCC-CNN BASED VOICE AUTHENTICATION SECURITY,” international journal of engineering technology and management sciences, pp. 358–363, Jul. 2022, doi: 10.46647/ijetms.2022.v06i04.0058.
N. Faisal Aljohani and E. Sami Jaha, “Visual Lip-Reading for Quranic Arabic Alphabets and Words Using Deep Learning,” Computer Systems Science and Engineering, vol. 46, no. 3, pp. 3037–3058, 2023, doi: 10.32604/csse.2023.037113.
F. Ahmad, S. Z. Yahya, Z. Saad, and A. R. Ahmad, “Tajweed classification using artificial neural network,” in 2018 International Conference on Smart Communications and Networking (SmartNets), IEEE, 2018, pp. 1–4.
G. Samara, E. Al-Daoud, N. Swerki, and D. Alzu’bi, “The recognition of holy qur’an reciters using the mfccs’ technique and deep learning,” Advances in Multimedia, vol. 2023, no. 1, p. 2642558, 2023.
Anwar, “Kesalahan Dalam Membaca Al-Qur’an.”
N. F. Aljohani and E. S. Jaha, “Visual Lip-Reading for Quranic Arabic Alphabets and Words Using Deep Learning.,” Computer Systems Science & Engineering, vol. 46, no. 3, 2023.
DOI: https://doi.org/10.47738/jads.v7i1.1062
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)