Adaptive Integration of Optuna Optimization and Stacking Ensemble Learning for Automated Work Competency Classification
Abstract
Artificial intelligence and machine learning are increasingly used to automate analytical and decision processes, including the evaluation of human competencies. However, traditional models often face challenges in accuracy and generalization when applied to linguistic data from interviews. This study aims to develop a model that integrates Optuna optimization and stacking ensemble learning to enhance the accuracy and interpretability of competency classification. Interview transcript data were processed using natural language processing techniques such as cleaning, tokenization, case folding, stopword removal, and stemming to ensure textual consistency. The text was then transformed into numerical representations using term frequency inverse document frequency weighting. To handle class imbalance, the synthetic minority oversampling technique was employed. Optuna was applied to optimize the hyperparameters of base models, including support vector classifier, Naïve Bayes, random forest, gradient boosting, and XGBoost. These optimized models were combined through a stacking ensemble to form the final classifier. The proposed model achieved an accuracy of 94 percent and a precision of 95 percent with macro and weighted F1 scores of 0.94. The results demonstrate stable and balanced performance across all competency categories, including analytical thinking, initiating action, problem solving, and work standards. Comparative analysis with previous studies in sentiment analysis, medical diagnosis, and financial forecasting confirmed that the integration of Optuna and stacking produces more robust and generalizable outcomes. The integration of Optuna optimization and stacking ensemble learning effectively improves classification performance while maintaining interpretability. The model demonstrates strong potential for automated competency evaluation in recruitment and human resource analytics. This framework can be extended to other linguistic datasets to support transparent and data-driven decision-making in artificial intelligence applications.
Keywords
Full Text:
PDFReferences
H. Qisong, W. Zixuan, Y. Wen, C. Minrui, C. Guoqiang, and Y. Yang, “Research on NOTAM Information Extraction of Civil Aviation with NLP,” 2023 IEEE 5th Int. Conf. Civ. Aviat. Saf. Inf. Technol., pp. 520–523, 2023, doi: 10.1109/iccasit58768.2023.10351768.
F. Zuo and H. Zhang, “College English Teaching Evaluation Model Using Natural Language Processing Technology and Neural Networks,” Mob. Inf. Syst., vol. 2022, 2022, doi: 10.1155/2022/7438464.
Rianto, A. B. Mutiara, E. P. Wibowo, and P. I. Santosa, “Improving the accuracy of text classification using stemming method, a case of non-formal Indonesian conversation,” J. Big Data, vol. 8, no. 1, pp. 1–16, 2021, doi: 10.1186/s40537-021-00413-1.
W. Wang, “Machine Learning-Based Intelligent Scoring of College English Teaching in the Field of Natural Language Processing,” Comput. Intell. Neurosci., vol. 2022, 2022, doi: 10.1155/2022/2754626.
X. Yue, “Research on the Semantic Analysis Method of Translation Corpus Based on Natural Language Processing,” Sci. Program., vol. 2022, 2022, doi: 10.1155/2022/3764230.
M. Novo-Lourés, R. Pavón, R. Laza, D. Ruano-Ordas, and J. R. Méndez, “Using natural language preprocessing architecture (NLPA) for big data text sources,” Sci. Program., vol. 2020, 2020, doi: 10.1155/2020/2390941.
M. Zhang, “Applications of Deep Learning in News Text Classification,” Sci. Program., vol. 2021, 2021, doi: 10.1155/2021/6095354.
A. Nilofer and S. Sasikala, “Models for High Dimensional Streaming Data,” IAENG Int. J. Comput. Sci., vol. 52, no. 12, pp. 4678–4691, 2025.
J. Atwan, M. Wedyan, Q. Bsoul, A. Hammadeen, and R. Alturki, “The Use of Stemming in the Arabic Text and Its Impact on the Accuracy of Classification,” Sci. Program., vol. 2021, 2021, doi: 10.1155/2021/1367210.
G. Sarin, P. Kumar, and M. Mukund, “Text classification using deep learning techniques: a bibliometric analysis and future research directions,” Benchmarking, 2023, doi: 10.1108/BIJ-07-2022-0454.
R. Patil, S. Boit, V. Gudivada, and J. Nandigam, “A Survey of Text Representation and Embedding Techniques in NLP,” IEEE Access, vol. 11, no. April, pp. 36120–36146, 2023, doi: 10.1109/ACCESS.2023.3266377.
V. Dogra et al., “A Complete Process of Text Classification System Using State-of-the-Art NLP Models,” Comput. Intell. Neurosci., vol. 2022, 2022, doi: 10.1155/2022/1883698.
R. Purushothaman, S. P. Rajagopalan, and C. Saravanakumar, “Efficient Analysis for Extracting Feature and Evaluation of Text Mining using Natural Language Processing Model,” Proc. 2021 IEEE Int. Conf. Innov. Comput. Intell. Commun. Smart Electr. Syst. ICSES 2021, pp. 1–5, 2021, doi: 10.1109/ICSES52305.2021.9633883.
P. Wang et al., “Classification of Proactive Personality: Text Mining Based on Weibo Text and Short-Answer Questions Text,” IEEE Access, vol. 8, pp. 97370–97382, 2020, doi: 10.1109/ACCESS.2020.2995905.
H. Wang and D. Zeng, “Fusing Logical Relationship Information of Text in Neural Network for Text Classification,” Math. Probl. Eng., vol. 2020, 2020, doi: 10.1155/2020/5426795.
V. S. Pendyala, N. Atrey, T. Aggarwal, and S. Goyal, “Enhanced Algorithmic Job Matching based on a Comprehensive Candidate Profile using NLP and Machine Learning,” Proc. - IEEE 8th Int. Conf. Big Data Comput. Serv. Appl. BigDataService 2022, pp. 183–184, 2022, doi: 10.1109/BigDataService55688.2022.00040.
C. G. Harris, “Age Bias: A Tremendous Challenge for Algorithms in the Job Candidate Screening Process,” Int. Symp. Technol. Soc. Proc., vol. 2022-Novem, pp. 1–5, 2022, doi: 10.1109/ISTAS55053.2022.10227135.
M. M. Hijazi, A. Zeki, and A. Ismail, “Arabic text classification: A review study on feature selection methods,” 2021 22nd Int. Arab Conf. Inf. Technol. ACIT 2021, pp. 1–6, 2021, doi: 10.1109/ACIT53391.2021.9677185.
S. M. Alsubhi, A. M. Alhothali, and A. A. Almansour, “AraBig5: The Big Five Personality Traits Prediction Using Machine Learning Algorithm on Arabic Tweets,” IEEE Access, vol. 11, no. June, pp. 112526–112534, 2023, doi: 10.1109/ACCESS.2023.3297981.
J. D’Souza, V. Kadam, P. Shinde, and K. Saxena, “The Quest for Fairness: A Comparative Study of Accuracy in AI Hiring Systems,” 2023 3rd Asian Conf. Innov. Technol. ASIANCON 2023, pp. 1–6, 2023, doi: 10.1109/ASIANCON58793.2023.10269895.
R. Moraes, L. L. Pinto, M. Pilankar, and P. Rane, “Personality Assessment Using Social Media for Hiring Candidates,” 2020 3rd Int. Conf. Commun. Syst. Comput. IT Appl. CSCITA 2020 - Proc., pp. 192–197, 2023, doi: 10.1109/CSCITA47329.2020.9137818.
I. Gupta, M. Jain, and P. Johri, “Smart-Hire Personality Prediction Using ML,” 2023 Int. Conf. Disruptive Technol. ICDT 2023, pp. 381–385, 2023, doi: 10.1109/ICDT57929.2023.10151367.
T. Madhu Midhan, P. Selvaraj, M. Harshavardan Kumar Raju, M. Bhanu Prakash Reddy, and T. Bhaskar, “Classification of Mental Health and Emotion of Human from Text using Machine Learning Approaches,” 2023 6th Int. Conf. Inf. Syst. Comput. Networks, ISCON 2023, pp. 1–7, 2023, doi: 10.1109/ISCON57294.2023.10111973.
T. Georgieva-Trifonova, “Research on Filtering Feature Selection Methods for E-Mail Spam Detection by Applying K-NN Classifier,” HORA 2022 - 4th Int. Congr. Human-Computer Interact. Optim. Robot. Appl. Proc., pp. 1–4, 2022, doi: 10.1109/HORA55278.2022.9799999.
F. Rollo, G. Bonisoli, and L. Po, “A Comparative Analysis of Word Embeddings Techniques for Italian News Categorization,” IEEE Access, vol. 12, pp. 25536–25552, 2024, doi: 10.1109/ACCESS.2024.3367246.
L. Jin and L. Zhang, “De-redundancy Relative Discrimination Criterion-based Feature Selection for Text Data,” Proc. Int. Jt. Conf. Neural Networks, vol. 2022-July, pp. 1–8, 2022, doi: 10.1109/IJCNN55064.2022.9892781.
K. Maheswar Reddy and R. Thandaiah Prabu, “Machine Learning Approach for Personality Prediction from Resume using XGBoost Classifier and Comparing with Novel Random Forest Algorithm to Improve Accuracy,” Proc. 8th IEEE Int. Conf. Sci. Technol. Eng. Math. ICONSTEM 2023, pp. 1–7, 2023, doi: 10.1109/ICONSTEM56934.2023.10142919.
S. A. Alshalif et al., “Alternative Relative Discrimination Criterion Feature Ranking Technique for Text Classification,” IEEE Access, vol. 11, no. July, pp. 71739–71755, 2023, doi: 10.1109/ACCESS.2023.3294563.
X. Tian, R. Pavur, H. Han, and L. Zhang, “A machine learning-based human resources recruitment system for business process management: using LSA, BERT and SVM,” Bus. Process Manag. J., vol. 29, no. 1, pp. 202–222, 2022, doi: 10.1108/BPMJ-08-2022-0389.
A. Saleem Raja, S. Balasubaramanian, P. Ganesan, J. Rajasekaran, and R. Karthikeyan, “Weighted ensemble classifier for malicious link detection using natural language processing,” Int. J. Pervasive Comput. Commun., 2023, doi: 10.1108/IJPCC-09-2022-0312.
C. Sun and B. Luo, “Analysis of English Writing Text Features Based on Random Forest and Logistic Regression Classification Algorithm,” Mob. Inf. Syst., vol. 2022, 2022, doi: 10.1155/2022/6306025.
F. Rabbi, S. Raut, N. Ullah, and I. Hossain, “Cardiovascular Risk Prediction Through Stacking Classifier †,” Enginering Proc., 2024.
M. Ali, M. Nasim, S. Anwar, M. Ali, and M. Nasim, “Stacking Random Forest Forest functioning functioning as as a a Meta Meta Stacking Classifier Classifier with with Random Classifier for Diabetes Diabetes Diseases Diseases Classification Classification Classifier for,” Procedia Comput. Sci., vol. 207, pp. 3459–3468, 2022, doi: 10.1016/j.procs.2022.09.404.
T. Ade, V. Ariandi, and S. Defit, “Enhancing Accuracy by Using Boosting and Stacking Techniques on the Random Forest Algorithm on Data from Social Media X,” vol. 16, no. 2, pp. 184–189, 2024.
I. R. Munthe, B. H. Rambe, F. Hanum, and A. T. Amanda, “Implementation of Stacking Technique Combining Machine Learning and Deep Learning Algorithms Using SMOTE to Improve Stock Market Prediction Accuracy,” vol. 5, no. 4, pp. 2079–2091, 2024.
DOI: https://doi.org/10.47738/jads.v7i2.1228
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)