Leveraging K-Nearest Neighbors with SMOTE and Boosting Techniques for Data Imbalance and Accuracy Improvement

Adyanata Lubis, Yuda Irawan, Junadhi Junadhi, Sarjon Defit

Abstract


This research addresses the issue of low accuracy in sentiment analysis on Israeli products on social media, initially achieving only 64% using the K-NN algorithm. Given the ongoing Israeli-Palestinian conflict, which has garnered widespread international attention and strong opinions, understanding public sentiment towards Israeli products is crucial. To improve accuracy, the study employs SMOTE to handle data imbalance and combines K-NN with boosting algorithms like AdaBoost and XGBoost, which were selected for their effectiveness in improving model performance on imbalanced and complex datasets. AdaBoost was chosen for its ability to enhance model accuracy by focusing on misclassified instances, while XGBoost was selected for its efficiency and robustness in handling large datasets with multiple features. The research process includes data pre-processing (cleaning, normalization, tokenization, stopwords removal, and stemming), labeling using a Lexicon-Based approach, and feature extraction with CountVectorizer and TF-IDF. SMOTE was applied to oversample the minority class to match the number of instances in the majority class, ensuring balanced representation before model training. A total of 1,145 datasets were divided into training and testing data with a ratio of 70:30. Results demonstrate that SMOTE increased K-NN accuracy to 77%. Interestingly, combining K-NN with AdaBoost after SMOTE achieved 72% accuracy, which, although lower than the 77% achieved with SMOTE alone, was higher than the 68% accuracy without SMOTE. This discrepancy can be attributed to the added complexity introduced by AdaBoost, which may not synergize as effectively with SMOTE as XGBoost does, particularly in this dataset's context. In contrast, K-NN with XGBoost after SMOTE reached the highest accuracy of 88%, demonstrating a more effective combination. Boosting without SMOTE resulted in lower accuracies: 68% for KNN+AdaBoost and 64% for KNN+XGBoost. The combination of K-NN with SMOTE and XGBoost significantly improves model accuracy and reliability for sentiment analysis on social media.

Keywords


K-NN, XGBoost, AdaBoost, SMOTE, Machine Learning

Full Text:

PDF

References


P. Wang, H. Shi, X. Wu, and L. Jiao, “Sentiment analysis of rumor spread amid covid-19: Based on weibo text,” Healthc., vol. 9, no. 10, 2021, doi: 10.3390/healthcare9101275.

A. Pamuji, “Performance of the K-Nearest Neighbors Method on Analysis of Social Media Sentiment,” Juisi, vol. 07, no. 01, pp. 32–37, 2021.

J. Mantik et al., “Application Of N-Gram On K-Nearest Neighbor Algorithm To Sentiment Analysis Of TikTok Shop Shopping Features,” J. Mantik, vol. 6, no. 3, pp. 2685–4236, 2022.

A. J. Mohammed, “Improving Classification Performance for a Novel Imbalanced Medical Dataset using SMOTE Method,” Int. J. Adv. Trends Comput. Sci. Eng., vol. 9, no. 3, pp. 3161–3172, 2020, doi: 10.30534/ijatcse/2020/104932020.

R. Natras, B. Soja, and M. Schmidt, “Ensemble Machine Learning of Random Forest, AdaBoost and XGBoost for Vertical Total Electron Content Forecasting,” Remote Sens., vol. 14, no. 15, pp. 1–34, 2022, doi: 10.3390/rs14153547.

S. Ghosal and A. Jain, “Depression and Suicide Risk Detection on Social Media using fastText Embedding and XGBoost Classifier,” Procedia Comput. Sci., vol. 218, pp. 1631–1639, 2022, doi: 10.1016/j.procs.2023.01.141.

R. Govindarajan, V. Balaji, J. Arumugam, T. A. Assegie, and R. Mothukuri, “Evaluation of sequential feature selection in improving the K-nearest neighbor classifier for diabetes prediction,” IAES Int. J. Artif. Intell., vol. 13, no. 2, pp. 1567–1573, 2024, doi: 10.11591/ijai.v13.i2.pp1567-1573.

A. Zamsuri, S. Defit, and G. W. Nurcahyo, “Classification Of Multiple Emotions In Indonesian Text Using The K-Nearest Neighbor Method,” J. Appl. Eng. Technol. Sci., vol. 4, no. 2, pp. 1012–1021, 2023, doi: 10.37385/jaets.v4i2.1964.

Nanda Ihwani Saputri, Yuliant Sibaroni, and Sri Suryani Prasetiyowati, “Covid-19 Fake News Detection on Twitter Based on Author Credibility Using Information Gain and KNN MethodsCovid-19 Fake News Detection on Twitter Based on Author Credibility Using Information Gain and KNN Methods,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 7, no. 1, pp. 185–192, 2023, doi: 10.29207/resti.v7i1.4871.

A. M. Elmogy, U. Tariq, A. Ibrahim, and A. Mohammed, “Fake Reviews Detection using Supervised Machine Learning,” Int. J. Adv. Comput. Sci. Appl., vol. 12, no. 1, pp. 601–606, 2021, doi: 10.14569/IJACSA.2021.0120169.

A. J. Barid, Hadiyanto, and A. Wibowo, “Optimization of the algorithms use ensemble and synthetic minority oversampling technique for air quality classification,” Indones. J. Electr. Eng. Comput. Sci., vol. 33, no. 3, pp. 1632–1640, 2024, doi: 10.11591/ijeecs.v33.i3.pp1632-1640.

A. N. Kasanah, M. Muladi, and U. Pujianto, “Penerapan Teknik SMOTE untuk Mengatasi Imbalance Class dalam Klasifikasi Objektivitas Berita Online Menggunakan Algoritma KNN,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 3, no. 2, pp. 196–201, 2019, doi: 10.29207/resti.v3i2.945.

C. Supriyanto, F. A. Rafrastara, A. Amiral, and ..., “Malware Detection Using K-Nearest Neighbor Algorithm and Feature Selection,” J. Media Inform. Budidarma, vol. 8, pp. 412–420, 2024, doi: 10.30865/mib.v8i1.6970.

S. Tomar, D. Dembla, and Y. Chaba, “Analysis and Enhancement of Prediction of Cardiovascular Disease Diagnosis using Machine Learning Models SVM, SGD, and XGBoost,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 4, pp. 469–479, 2024, doi: 10.14569/IJACSA.2024.0150449.

A. Mishra, S. Mishra, and P. Jain, “Malware Category Prediction Using KNN And SVM Classifiers,” Int. J. Mech. Eng. Technol., vol. 10, no. 02, pp. 787–797, 2019.

W. Chimphlee and S. Chimphlee, “Hyperparameters optimization XGBoost for network intrusion detection using CSE-CIC-IDS 2018 dataset,” IAES Int. J. Artif. Intell., vol. 13, no. 1, pp. 817–826, 2024, doi: 10.11591/ijai.v13.i1.pp817-826.

Nur Ghaniaviyanto Ramadhan, “Indonesian Online News Topics Classification using Word2Vec and K-Nearest Neighbor,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 5, no. 6, pp. 1083–1089, 2021, doi: 10.29207/resti.v5i6.3547.

D. Daimari, S. Mondal, B. Brahma, and A. Nag, “Favorite Book Prediction System Using Machine Learning Algorithms,” J. Appl. Eng. Technol. Sci., vol. 4, no. 2, pp. 983–991, 2023, doi: 10.37385/jaets.v4i2.1925.

M. A. Rosid, A. S. Fitrani, I. R. I. Astutik, N. I. Mulloh, and H. A. Gozali, “Improving Text Preprocessing for Student Complaint Document Classification Using Sastrawi,” IOP Conf. Ser. Mater. Sci. Eng., vol. 874, no. 1, pp. 0–6, 2020, doi: 10.1088/1757-899X/874/1/012017.

Y. HaCohen-Kerner, D. Miller, and Y. Yigal, “The influence of preprocessing on text classification using a bag-of-words representation,” PLoS One, vol. 15, no. 5, pp. 1–22, 2020, doi: 10.1371/journal.pone.0232525.

M. Novo-Lourés, R. Pavón, R. Laza, D. Ruano-Ordas, and J. R. Méndez, “Using natural language preprocessing architecture (NLPA) for big data text sources,” Sci. Program., vol. 2020, 2020, doi: 10.1155/2020/2390941.

W. Bourequat and H. Mourad, “Sentiment Analysis Approach for Analyzing iPhone Release using Support Vector Machine,” Int. J. Adv. Data Inf. Syst., vol. 2, no. 1, pp. 36–44, 2021, doi: 10.25008/ijadis.v2i1.1216.

S. K. Dirjen et al., “Terakreditasi SINTA Peringkat 2 Analisis Pengaruh Data Scaling Terhadap Performa Algoritme Machine Learning untuk Identifikasi Tanaman,” J. Resti, vol. 4, no. 1, pp. 117–122, 2020.

N. Garg and K. Sharma, “Text pre-processing of multilingual for sentiment analysis based on social network data,” Int. J. Electr. Comput. Eng., vol. 12, no. 1, pp. 776–784, 2022, doi: 10.11591/ijece.v12i1.pp776-784.

L. Hickman, S. Thapa, L. Tay, M. Cao, and P. Srinivasan, “Text Preprocessing for Text Mining in Organizational Research: Review and Recommendations,” Organ. Res. Methods, vol. 25, no. 1, pp. 114–146, 2022, doi: 10.1177/1094428120971683.

R. Kalaivani and R. Marivendan, “The effect of stop word removal and stemming in datapreprocessing,” Ann. R.S.C.B, vol. 25, no. 6, pp. 739–746, 2021.

A. N. Ulfah, M. K. Anam, N. Y. Sidratul Munti, S. Yaakub, and M. B. Firdaus, “Sentiment Analysis of the Convict Assimilation Program on Handling Covid-19,” JUITA J. Inform., vol. 10, no. 2, p. 209, 2022, doi: 10.30595/juita.v10i2.12308.

Junadhi, Agustin, M. Rifqi, and M. K. Anam, “Sentiment Analysis of Online Lectures using K-Nearest Neighbors based on Feature Selection,” J. Nas. Pendidik. Tek. Inform., vol. 11, no. 3, pp. 216–225, 2022, doi: 10.23887/janapati.v11i3.51531.

T. Colibazzi et al., “Identifying Splitting Through Sentiment Analysis,” J. Pers. Disord., vol. 37, no. 1, pp. 36–48, 2023, doi: 10.1521/pedi.2023.37.1.36.

M. Rezwanul, A. Ali, and A. Rahman, “Sentiment Analysis on Twitter Data using KNN and SVM,” Int. J. Adv. Comput. Sci. Appl., vol. 8, no. 6, pp. 19–25, 2017, doi: 10.14569/ijacsa.2017.080603.

S. Wang, Y. Dai, J. Shen, and J. Xuan, “Research on expansion and classification of imbalanced data based on SMOTE algorithm,” Sci. Rep., vol. 11, no. 1, pp. 1–11, 2021, doi: 10.1038/s41598-021-03430-5.

Y. Ding, H. Zhu, R. Chen, and R. Li, “An Efficient AdaBoost Algorithm with the Multiple Thresholds Classification,” Appl. Sci., vol. 12, no. 12, 2022, doi: 10.3390/app12125872.

J. Xu, Y. Jiang, and C. Yang, “Landslide Displacement Prediction during the Sliding Process Using XGBoost, SVR and RNNs,” Appl. Sci., vol. 12, no. 12, 2022, doi: 10.3390/app12126056




DOI: https://doi.org/10.47738/jads.v5i4.343

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0