Dispute on Security Framework Model of MFCC Mixed Methods in Speech Recognition System 

Heni Ispur Pratiwi, Iman Herwidiana Kartowisastro, Benfano Soewito, Widodo Budiharto

Abstract


An audio recording system device has unprecedented activities of its authorized users which in a particular way cause vulnerability to the system. It starts to get into a fuzzy condition and deteriorate the system sensitivity in detecting unauthorized access to pass through, then the system inclination may occur. One case is when separate users picked speech voices with similar keywords to set their usernames or password. Moreover, when users are siblings or twins that could have merely similar voices.  Troublesome of this situation leads to a less sensitive manner of a security system, and in some situations, the system could operate blocking authorized users themselves to get access. This paper defines a proposed method to resolve the situation by combining Mel Frequency Cepstral Coefficient with other methodologies, which have been implemented for many other research’ specific objectives as well. This paper displays to prove its combination with an interval scoring in Fuzzy Relation complements a resolution to tackle the security of fuzzy issues mentioned. The Mel Scale has its capacity of delivering extractions output from audio input data, it is called as spectral centroids which refer to humans’ voices or an individual's voice features. Some spectral centroids get merely similar results due to those inclinations mentioned. This paper exposes Fuzzy Relation method to fit the need of verification procedures thorough its interval scale on any fuzzy features. The objective of verification procedure is to gain consistency measured scales, and security warrant remains valid. The inhouse experiments served to give user A of [0.49, 1.18] interval, user B of [0.76,1.07] interval, and user C of [0.44,0.95] interval, and those interval numbers are proposed to cap other login users accounts unto theirs.


Keywords


Spectral Centroid; Mel Scale; Fourier Transform; Fuzzy; Speech Character

Full Text:

PDF

References


H.I. Pratiwi, I.H. Kartowisastro, B. Soewito, and W. Budiharto, “Short Time Fourier Transform in Reinvigorating Distinctive Facts of Individual Spectral Centroid of Mel Frequency Numeric for Security Authentication,” International Journal of Innovative Computing, Information and Control (IJICIC), Vol. 20, February 2024.

D.R. Kisku, et al., “Advances in Biometric for Secure Human Authentication and Recognition,” CRC Press, Taylor and Francis Group, 2014.

H. Gyulyustan, H., and S. Enkov, “Experimental Speech Recognition System Based on Raspberry Pi 3,” IOSR Journal of Computer Engineering (IIOSR - JCE)), Volume 19, Issue 3, p. 107-112, 2017.

S.I. Khamlich, Atouf, F. Khamlich, and M. Benrabh, “Performance Evaluation and Implementations of MFCC, SVM and MLP Algorithms in the FPGA Board,” International Journal of Electrical and Computer Engineering Systems Vol.12 No. 3, 2021.

H.C. Shantakumar, G.S. Nagaraja, and M. Basthikodi, “Performance Evolution of Face and Speech Recognition system using DTCWT and MFCC Features,” Turkish Journal of Computer and Mathematics Education Vol.112 No. 3, 395 – 404, 2021.

D.Y. Mohammed, K. Al-Karawi, and A. Aljuboori, “Robust speaker verifications by combining MFCC and Entrocy in noisy conditions,” Bulletin of Electrical Engineering and Informatics, vol.10, no.4, pp.2310-2319, 2021.

B. Birch, C.A. Griffith, and A. Morgan, “Environmental Effects on Reliability and Accuracy of MFCC Based voice Recognition for Industrial Human-Robot-Interaction,” Proc IMechE Part B: J Engineering Manufacture, Vol. 235(12) 1939–1948, ImechE, 2021.

V.S. Wicaksana, and A. Zahra, “Spoken Language Identification on Local Language Using MFCC, Random Forest, KNN, and GMM,” IJACSA Int. Journal of advanced of comp science and applications Vol.12 No.5, 2021.

M. Maseri, and M. Mamat, “Performance Analysis of Implemented MFCC and HMM-based Speech Recognition System,” Authorized licensed of © IEEE UNIVERSITY SABAH MALAYSIA IEEE Xplore, 2020, Downloaded on November 03, 2022 at 15:58:39 UTC.

A. Chowdhury, and A. Ross, “Fusing MFCC and LPC Features Using 1D Triplet CNN for Speaker Recognition in Severely Degraded Audio Signals,” IEEE Transactions on Information Forensic and Security, Vol. 15, 2020.

M. Malik, M.K.Malik, K.Mehmood, and I. Makhdoom, “Automatic Speech Recognition: A Survey,” Springer Science+Business Media, LLC, part of Springer Nature 2020.

T. Gunendradasan, B.Wickramasinghe, P. Ngoc Le, E. Ambikairajah, and J. Epps, “Detection of Replay-Spoofing Attacks Using Frequency Modulation Features,” Hyderabad: Interspeech 2018.

X. Li and M. Mills, “Vocal Features: From Voice Identification to Speech Recognition by Machine,” Technology and Culture, Vol. 60, No. 2, Johns Hopkins University Press, 2019.

Y. Chen, X. Yuan, A. Wang, K. Chen, S. Zhang, and H. Huang, “Manipulating Users’ Trust on Amazon Echo: Compromising Smart Home from Outside,” EAI Endorsed Transactions on Security and Safety 2020, http://creativecommons.org/licenses/by/3.0/.

J.S. Edu, J.M. Such, and G. Suarez-Tangli, “Smart Home Personal Assistants: A Security and Privacy Review,” ACM Computer Survey 1, August 2020.

A. Fazel, W. Yang, Y. Liu, R. Barra-Chicote, Y. Meng, R. Maas, and J. Droppo, “SynthASR: Unlocking Synthetic Data for Speech Recognition,” ArXiv: 2106.07803v1 [cs/LG] 14 Jun 2021.

H. Voss, H. Wersing, and S. Kopp, “Addressing Data Scarcity in Multimodal User State Recognition by Combining Semi-Supervised and Supervised Learning,” ICMI ’21 Companion, October 18–22, 2021, Montréal, QC, Canada, © 2021 Copyright held by the owner/author(s). Publication rights license ACM ISBN 978-1-4503- 8471-1/21/10.

J.L.K.E. Fendji, “Automatic Speech Recognition and Limited Vocabulary: A Survey,” © Elsevier, 2022

H.I. Pratiwi, I.H. Kartowisastro, B. Soewito, and W. Budiharto, “Adopting Centroid and Bandwidth to Shape Security Line,” IEEE Xplore Conference Icosnikom, Medan, Indonesia 2022.

T. Safavi, “Automatic Speaker, Age-group and Gender Identification from Children’s Speech, Computer Speech & Language,” doi: 10.1016/j.csl.2018.01.001.

H. Sujadi, “Sistem Pengolahan Suara Menggunakan Agoritma FFT (Fast Fourier Transform),” Proceeding SINTAK ISBN: 978-602-8557-20-7101, 2017.

S. Meignen, D.H Pham, M.A. Colominas, “On the Use of Short-Time Fourier Transform and Synchrosqueezing-Based Demodulation for the Retrieval of the Modes of Multicomponent Signals,” ©2020 published by Elsevier, Available online: https://www.elsevier.com/open-access/userlicense/1.0/.

W. Mustikarini, R. Hidayat, and A. Bejo, “Real-Time Indonesian Language Speech Recognition with MFCC Algorithms and Python - Based SVM,” IJITEE Vol.3 , No.2, 2019.

A.H. Nour-Eldin, Mel-Frequency Cepstral Coefficient-Based Bandwidth Extension of Narrowband Speech, Interspeech Brisbane, Australia, Copyright © 2008 ISCA.

T. Furoh, F. Takahiro, M. Nakayama, and n. Takanobu, “A study of degraded-speech identification based on spectral centroid,” Inter-noise , I-INCE, Classification of Subjects Number(s): 01.4, 2014.

Hua-Peng Zhang, “On the Construction of Fuzzy Betweenness Relations from Metrics,” Fuzzy Set and System, 390, 118 - 137 ScienceDirect © 2020 published by Elsevier.

J.Y. Choi, “Varying acoustic-phonemic ambiguity reveals that talker normalization is obligatory in speech processing,” Published online: 7 February 2018, © The Psychonomic Society, Inc. 2018.

C. Gowrishankar, “Properties of Composition of Fuzzy Relations and its Verifications,” International Journal of Management and Humanities (IJMH) ISSN: 2394 (online), Vol.4 Issue-5, January 2020.

J.H.L. Hansena and H. Borila, “On the Issues of Intra-Speaker Variability and Realism in Speech, Speaker, and Language Recognition Tasks,” © 2018 published by Elsevier.




DOI: https://doi.org/10.47738/jads.v6i3.689

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0