Implementation of Stacking Technique Combining Machine Learning and Deep Learning Algorithms Using SMOTE to Improve Stock Market Prediction Accuracy
Abstract
This study introduces a stacking technique that integrates machine learning (ML) and deep learning (DL) algorithms to enhance the accuracy of stock market trend predictions. The stacking model utilizes XGBoost and Random Forest as base models from the ML domain, while Logistic Regression and LSTM (Long Short-Term Memory) function as meta models to optimize predictive accuracy. A significant challenge in stock market data is class imbalance, where certain trends, such as stock price drops, are underrepresented. To mitigate this, we applied the Synthetic Minority Over-sampling Technique (SMOTE) to generate synthetic data for the minority class. This approach helps the model better capture patterns from the underrepresented data while preserving essential information from the majority class. The implementation of SMOTE, coupled with the stacking technique, yielded a substantial improvement in prediction accuracy. The results showed that the Random Forest algorithm achieved an accuracy of 85% with precision, recall, and F1-score all at 85%, while XGBoost and Logistic Regression achieved accuracies of 82% and 81% respectively. For the deep learning models, LSTM reached an accuracy of 83%, while the Stacking Meta Model with LSTM achieved an accuracy of 83% with slightly better precision and recall at 84%. The stacking model, with Logistic Regression as the meta model, ultimately achieved the highest accuracy of 86%, outperforming individual models such as SVM (Support Vector Machine), LSTM, Random Forest, and Logistic Regression (LR). These findings demonstrate the efficacy of combining SMOTE with stacking to address data imbalance and improve stock market predictions. The novelty of this study lies in the integration of advanced ML and DL models within a stacking framework to handle class imbalance in financial datasets. Future research will explore the deployment of this model in a real-time web-based application to support investor decision-making in stock market trend analysis.
Keywords
Full Text:
PDFReferences
U. Septepana, P. L. Obrien, A. F. Surbakti, F. Astuty, and W. A. Ginting, “Analysis of Macroeconomic Factors that Influence the Movement of the Composite Stock Price Index (IHSG) on the Indonesian Stock Exchange for the Period 2013-2023,” IJEBIR, vol. 3, no. 4, pp. 142–155, 2024, [Online]. Available: www.idxchannel.com
Fuad and I. Yuliadi, “Determinants of the Composite Stock Price Index (IHSG) on the Indonesia Stock Exchange,” Journal of Economics Research and Social Sciences, vol. 5, no. 1, pp. 27–41, 2021, doi: 10.18196/jerss.v5i1.11002.
C. Kumala Dewi and R. Masithoh Haryadi, “Volatility Analysis Of Indonesia Composite Index During COVID-19,” JEM, vol. 15, no. 2, pp. 138–148, 2021, doi: 10.30650/jem.v15i2.2225.
A. Logitama, L. Setiawan, and A. Hayat, “Control Behaviors Affecting Investors Investment Decision Making (Studies on Students at Higher Education South Kalimantan),” Business, and Accounting Research (IJEBAR) Peer Reviewed-International Journal, vol. 5, no. 1, pp. 278–291, 2021, [Online]. Available: https://jurnal.stie-aas.ac.id/index.php/IJEBAR
A. S. Rakhmat and M. H. Fahamsyah, “Literature Review on Detecting Bubbles in Stock Market,” East Asian Journal of Multidisciplinary Research, vol. 2, no. 4, pp. 1499–1504, Apr. 2023, doi: 10.55927/eajmr.v2i4.3828.
H. Soegoto and I. Afiansyah, “Young People Play Stocks Without Fear of Cut Loss,” Indonesian Community Service and Empowerment Journal (IComSE), vol. 4, no. 1, pp. 313–317, 2023, doi: 10.34010/icomse.v4i1.7914.
A. Anwar, A. A. Patil, S. Choudari, Chetna, and C. S. Kiran, “Machine Learning Insights for Stock Market Trend Identification,” International Journal of Intelligent Systems and Applications in Engineering IJISAE, vol. 12, no. 10s, pp. 567–575, 2024, [Online]. Available: https://orcid.org/0000-0002-6756-6514
S. L. V. V. D. Sarma, D. V. Sekhar, and G. Murali, “Stock market analysis with the usage of machine learning and deep learning algorithms,” Bulletin of Electrical Engineering and Informatics, vol. 12, no. 1, pp. 552–560, Feb. 2023, doi: 10.11591/eei.v12i1.4305.
S. Kompella and K. C. Chilukuri, “Stock Market Prediction Using Machine Learning Methods,” International Journal of Computer Engineering and Technology, vol. 10, no. 3, pp. 20–30, 2019, [Online]. Available: www.jifactor.com
N. G. Ramadhan, “Comparative Analysis of ADASYN-SVM and SMOTE-SVM Methods on the Detection of Type 2 Diabetes Mellitus,” Scientific Journal of Informatics, vol. 8, no. 2, pp. 276–282, Nov. 2021, doi: 10.15294/sji.v8i2.32484.
X. T. Dang and T. T. Le, “KNN-SMOTE: An Innovative Resampling Technique Enhancing the Efficacy of Imbalanced Biomedical Classification,” in Machine Learning and Other Soft Computing Techniques: Biomedical and Related Applications, N. Hoang Phuong, N. T. Huyen Chau, and V. Kreinovich, Eds., Cham: Springer Nature Switzerland, 2024, pp. 111–121. doi: 10.1007/978-3-031-63929-6_11.
P. P. Putra, M. K. Anam, S. Defit, and A. Yunianta, “Enhancing the Decision Tree Algorithm to Improve Performance Across Various Datasets,” INTENSIF: Jurnal Ilmiah Penelitian dan Penerapan Teknologi Sistem Informasi, vol. 8, no. 2, pp. 200–212, Aug. 2024, doi: 10.29407/intensif.v8i2.22280.
V. Talasila, M. V Mohan, and N. M. R, “Enhancing Text-to-Image Synthesis with an Improved Semi-Supervised Image Generation Model Incorporating N-Gram, Enhanced TF-IDF, and BOW Techniques,” International Journal of Intelligent Systems and Applications in Engineering , vol. 11, no. 7s, pp. 381–397, 2023, [Online]. Available: www.ijisae.org
D. Tarwidi, S. R. Pudjaprasetya, D. Adytia, and M. Apri, “An optimized XGBoost-based machine learning method for predicting wave run-up on a sloping beach,” MethodsX, vol. 10, pp. 1–12, Aug. 2023, doi: 10.1145/2939672.2939785.
L. K. Shrivastav and R. Kumar, “An Ensemble of Random Forest Gradient Boosting Machine and Deep Learning Methods for Stock Price Prediction,” Journal of Information Technology Research, vol. 15, no. 1, pp. 1–19, Nov. 2021, doi: 10.4018/jitr.2022010102.
P. N. Anggreyani and W. Maharani, “Hoax Detection Tweets of the COVID-19 on Twitter Using LSTM-CNN with Word2Vec,” Jurnal Media Informatika Budidarma, vol. 6, no. 4, pp. 2432–2437, Oct. 2022, doi: 10.30865/mib.v6i4.4564.
M. Vilares Ferro, Y. Doval Mosquera, F. J. Ribadas Pena, and V. M. Darriba Bilbao, “Early stopping by correlating online indicators in neural networks,” Neural Networks, vol. 159, pp. 109–124, Feb. 2023, doi: 10.1016/j.neunet.2022.11.035.
B. L. V. S. R. Krishna, V. Mahalakshmi, and G. K. M. Nukala, “A Stacking Model for Outlier Prediction using Learning Approaches,” International Journal of Intelligent Systems and Applications in Engineering, vol. 12, no. 2s, pp. 629–638, 2023, [Online]. Available: https://www.kaggle.com/datasets/johnsmith88/heart-
A. Desfiandi and B. Soewito, “Student Graduation Time Prediction Using Logistic Regression, Decision Tree, Support Vector, And Adaboost Ensemble Learning,” International Journal of Information System and Computer Science) IJISCS, vol. 7, no. 3, pp. 195–199, 2023, doi: 10.56327/ijiscs.v7i2.1579.
L. K. Ramasamy, S. Kadry, Y. Nam, and M. N. Meqdad, “Performance analysis of sentiments in Twitter dataset using SVM models,” International Journal of Electrical and Computer Engineering, vol. 11, no. 3, pp. 2275–2284, Jun. 2021, doi: 10.11591/ijece.v11i3.pp2275-2284.
P. Bintoro, T. H. Andika, A. F. Yulia, and P. Widiandana, “Sentiment Analysis on Twitter Using Machine Learning Approach,” International Journal of Software Engineering and Informatics, vol. 1, no. 1, pp. 33–39, 2023, [Online]. Available: https://journal.aisyahuniversity.ac.id/index.php/IJosei
A. Romadhony, S. Al Faraby, R. Rismala, U. N. Wisesti, and A. Arifianto, “Sentiment Analysis on a Large Indonesian Product Review Dataset,” Journal of Information Systems Engineering and Business Intelligence, vol. 10, no. 1, pp. 167–178, 2024, doi: 10.20473/jisebi.10.1.167-178.
M. K. Anam, S. Defit, Haviluddin, L. Efrizoni, and M. B. Firdaus, “Early Stopping on CNN-LSTM Development to Improve Classification Performance,” Journal of Applied Data Sciences, vol. 5, no. 3, pp. 1175–1188, 2024, doi: 10.47738/jads.v5i3.312.
A. Muhaddisi, B. N. Prastowo, and D. U. K. Putri, “Sentiment Analysis With Sarcasm Detection On Politician’s Instagram,” IJCCS (Indonesian Journal of Computing and Cybernetics Systems), vol. 15, no. 4, p. 349, Oct. 2021, doi: 10.22146/ijccs.66375.
S. A. H. Bahtiar, C. K. Dewa, and A. Luthfi, “Comparison of Naïve Bayes and Logistic Regression in Sentiment Analysis on Marketplace Reviews Using Rating-Based Labeling,” Journal of Information Systems and Informatics, vol. 5, no. 3, pp. 915–927, Aug. 2023, doi: 10.51519/journalisi.v5i3.539.
DOI: https://doi.org/10.47738/jads.v5i4.421
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)