HU Variance Moment Optimizes Keyframe Selection Based on Deep Learning for Violence Detection

Sukmawati Anggraeni Putri, Pulung Nurtantio Andono, Purwanto Purwanto, Moch Arief Soeleman

Abstract


Violence in public spaces poses a serious threat to individuals and society. Manual monitoring and violence detection require much time and human resources, ultimately hindering detection accuracy and speed. Therefore, an automated method is needed to detect violence to ensure fast and efficient action. Along with technological advances, violence detection research has adopted various methods and models, including deep learning, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs). In this study, the classification process for detecting violence and non-violence uses the VGG19 model, one of the CNN models that has good performance with limited computing. In addition, the Long Short-Term Memory (LSTM) model is the best RNN model for processing temporal data in videos. However, this performance will decrease with noise and irrelevant data in the classification process. Therefore, to optimize deep learning performance, this study in the pre-processing phase selects keyframes in frame extraction using the Hu Variance Moment Technique. This method calculates each frame’s Hu and Variance Moment values and selects keyframes based on high Hu values. Next, we use Adaptive Moment Estimation (Adam) to optimize the gradient of the selected keyframes. This study produces a Hu19LSTM model tested on three datasets: hockey fight, crowd, and AIRTLab. The proposed Hu19LSTM model produces an accuracy of 97% on the Hockey Fight dataset, 97% on the Crowd dataset, and 95% on the AIRTLab dataset. These results indicate that the Hu19LSTM model can increase its accuracy on the hockey fight and Crowd dataset by 97%.

Keywords


Violence Detection; Hu Variance Moment Technique; VGG19; Long Short-Term Memory; Adaptive Moment Estimation

Full Text:

PDF

References


J. Feng, Y. Liang, and L. Li, “Anomaly Detection in Videos Using Two-Stream Autoencoder with Post Hoc Interpretability,” Comput. Intell. Neurosci., vol. 2021, 2021, doi: 10.1155/2021/7367870.

S. A. Velastin, B. A. Boghossian, and M. A. Vicencio-Silva, “A motion-based image processing system for detecting potentially dangerous situations in underground railway stations,” Transp. Res. Part C Emerg. Technol., vol. 14, no. 2, pp. 96–113, 2006, doi: 10.1016/j.trc.2006.05.006.

M. Y. Yang, W. Liao, Y. Cao, and B. Rosenhahn, “Video event recognition and anomaly detection by combining gaussian process and hierarchical Dirichlet process models,” Photogramm. Eng. Remote Sensing, vol. 84, no. 4, pp. 203–214, 2018, doi: 10.14358/PERS.84.4.203.

S. U. Khan, I. U. Haq, S. Rho, S. W. Baik, and M. Y. Lee, “Cover the violence: A novel deep-learning-based approach towards violence-detection in movies,” Appl. Sci., vol. 9, no. 22, 2019, doi: 10.3390/APP9224963.

A. Ben Mabrouk and E. Zagrouba, “Abnormal behavior recognition for intelligent video surveillance systems: a review,” Expert Syst. Appl., no. 91, pp. 480–491, 2018, doi: 10.1016/j.eswa.2017.09.029.

Z. Shao, J. Cai, and Z. Wang, “Smart Monitoring Cameras Driven Intelligent Processing to Big Surveillance Video Data,” IEEE Trans. Big Data, vol. 4, no. 1, pp. 105–116, 2017, doi: 10.1109/tbdata.2017.2715815.

Barbara Kitchenham, “Systematic Review in Software Engineering – Where We Are and Where We Should Be Going,” in EAST ’12: Proceedings of the 2nd international workshop on Evidential assessment of software technologies, 2007, doi: 10.1145/2372233.2372235.

M. Biswas et al., “State-of-the-Art Violence Detection Techniques: A review,” Asian J. Res. Comput. Sci., no. February, pp. 29–42, 2022, doi: 10.9734/ajrcos/2022/v13i130305.

I. Serrano, O. Deniz, J. L. Espinosa-Aranda, and G. Bueno, “Fight Recognition in Video Using Hough Forests and 2D Convolutional Neural Network,” IEEE Trans. Image Process., vol. 27, no. 10, pp. 4787–4797, 2018, doi: 10.1109/TIP.2018.2845742.

P. Contardo, P. Sernani, N. Falcionelli, and A. F. Dragoni, “Deep learning for law enforcement: A survey about three application domains,” CEUR Workshop Proc., vol. 2872, no. July, pp. 36–45, 2021.

M. Shoaib and N. Sayed, “A Deep Learning Based System for the Detection of Human Violence in Video Data,” Trait. du Signal, vol. 38, no. 6, pp. 1623–1635, 2021, doi: 10.18280/ts.380606.

U. Usman, F. Yunita, and M. R. Ridha, “Improving Classification Accuracy of Local Coconut Fruits with Image Augmentation and Deep Learning Algorithm Convolutional Neural Networks ( CNN ),” vol. 6, no. 1, pp. 1–19, 2025.

M. Patel, “Real-Time Violence Detection Using CNN-LSTM,” 2021.

R. Halder and R. Chatterjee, “CNN-BiLSTM Model for Violence Detection in Smart Surveillance,” SN Comput. Sci., vol. 1, no. 4, 2020, doi: 10.1007/s42979-020-00207-x.

H. M. Bin Jahlan, “Detecting Violence in Video Based on Deep Features Fusion Technique,” pp. 1–15.

F. U. M. Ullah, A. Ullah, K. Muhammad, I. U. Haq, and S. W. Baik, “Violence detection using spatiotemporal features with 3D convolutional neural network,” Sensors, vol. 19, no. 11, pp. 1–15, 2019, doi: 10.3390/s19112472.

T. Haque, F. F. Ahmed, S. M. I. Ahmed, and M. Siam, “Optical Flow based Violence Detection from Video Footage using Hybrid MobileNet and Bi-LSTM,” no. September, 2023.

S. Mukherjee, R. Saini, P. Kumar, P. P. Roy, D. P. Dogra, and B.-G. Kim, “Fight Detection in Hockey Videos using Deep Network,” J. Multimed. Inf. Syst., vol. 4, no. 4, pp. 225–232, 2017.

S. Sudhakaran and O. Lanz, “Learning to detect violent videos using convolutional long short-term memory,” 2017 14th IEEE Int. Conf. Adv. Video Signal Based Surveillance, AVSS 2017, 2017, doi: 10.1109/AVSS.2017.8078468.

Y. Sun, G. Wen, and J. Wang, “Weighted spectral features based on local Hu moments for speech emotion recognition,” Biomed. Signal Process. Control, vol. 18, pp. 80–90, 2015, doi: 10.1016/j.bspc.2014.10.008.

H. Ming-Kuei, “Visual pattern recognition by moment invariants,” IRE Trans. Inf. Theory, pp. 179–188, 1962.

S. Letchmunan, U. M. Butt, F. H. Hassan, S. Zia, and A. Baqir, “Detecting video surveillance using VGG19 convolutional neural networks,” Int. J. Adv. Comput. Sci. Appl., vol. 11, no. 2, pp. 674–682, 2020, doi: 10.14569/ijacsa.2020.0110285.

E. Bermejo Nievas, O. Deniz Suarez, G. Bueno García, and R. Sukthankar, “Violence Detection in Video Using Computer Vision Techniques,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 6855 LNCS, no. PART 2, pp. 332–339, 2011.

P. Sernani, N. Falcionelli, S. Tomassini, P. Contardo, and A. F. Dragoni, “Deep Learning for Automatic Violence Detection: Tests on the AIRTLab Dataset,” IEEE Access, vol. 9, pp. 160580–160595, 2021, doi: 10.1109/ACCESS.2021.3131315.

Irfanullah, T. Hussain, A. Iqbal, B. Yang, and A. Hussain, “Real time violence detection in surveillance videos using Convolutional Neural Networks,” Multimed. Tools Appl., vol. 81, no. 26, pp. 38151–38173, Nov. 2022, doi: 10.1007/s11042-022-13169-4.

A.-M. R. Abdali and A.-T. Rana F., “Robust Real-Time Violence Detection in Video Using CNN And LSTM,” 2019 2nd Sci. Conf. Comput. Sci., pp. 104–108, 2019.

K. Gkountakos, K. Ioannidis, T. Tsikrika, S. Vrochidis, and I. Kompatsiaris, “Crowd Violence Detection from Video Footage,” Proc. - Int. Work. Content-Based Multimed. Index., vol. 2021-June, no. September, 2021, doi: 10.1109/CBMI50038.2021.9461921.




DOI: https://doi.org/10.47738/jads.v6i2.648

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0