Assessing Large Language Models for Zero-Shot Dynamic Question Generation and Automated Leadership Competency Assessment
Abstract
Automated interview systems powered by artificial intelligence often rely on fine-tuned models and annotated datasets, limiting their adaptability to new leadership competency frameworks. Large language models have shown potential for generating questions and assessing answers, yet their zero-shot performance, operating without task-specific retraining remains underexplored in leadership assessment. This study examines the zero-shot capability of two models, Qwen 32B and GPT-4o-mini, within a multi-turn self-interview framework. Both models dynamically generated questions, interpreted responses, and assigned scores across ten leadership competencies. Professionals representing the role of Digital Marketing and Account Manager participated, each completing two AI-led interview sessions. Model outputs were evaluated by certified experts using a structured rubric across three dimensions: quality of behavioral insights, relevance of follow-up questions, and fit of assigned scores. Results indicate that Qwen 32B generated richer insights than GPT-4o-mini (mean = 2.86 vs. 2.62; p less than 0.01) and provided more differentiated assessments across competencies. GPT-4o-mini produced more consistent follow-up questions but lacked depth in interpretation, often yielding generic outputs. Both models struggled with accurate scoring of candidate responses, reflected in low answer score ratings (Qwen mean = 2.35; GPT mean = 2.21). These findings suggest a trade-off between insight richness and scoring stability, with both models demonstrating limited ability to fully capture nuanced leadership behaviors. This study offers one of the first empirical benchmarks of zero-shot model performance in leadership interviews. It underscores both the promise and current limitations of deploying such systems for scalable assessment. Future research should explore competency-specific prompt strategies, fairness evaluation across demographic groups, and domain-adapted fine-tuning to improve accuracy, reliability, and ethical alignment in high-stakes recruitment contexts.
Keywords
Full Text:
PDFReferences
J. S. Black and P. van Esch, “AI-enabled recruiting: What is it and how should a manager use it?,” Bus Horiz, vol. 63, no. 2, pp. 215–226, Mar. 2020, doi: 10.1016/j.bushor.2019.12.001.
D. W. Otter, J. R. Medina, and J. K. Kalita, “A Survey of the Usages of Deep Learning in Natural Language Processing,” Jul. 2018, [Online]. Available: http://arxiv.org/abs/1807.10854
S. Serrano, Z. Brumbaugh, and N. A. Smith, “Language Models: A Guide for the Perplexed,” Nov. 2023, [Online]. Available: http://arxiv.org/abs/2311.17301
M. F. Gonzalez et al., “Allying with AI? Reactions toward human-based, AI/ML-based, and augmented hiring processes,” Comput Human Behav, vol. 130, May 2022, doi: 10.1016/j.chb.2022.107179.
R. L. F. Garcia, Y. K. Huang, and L. Kwok, “Virtual interviews vs. LinkedIn profiles: Effects on human resource managers’ initial hiring decisions,” Tour Manag, vol. 94, Feb. 2023, doi: 10.1016/j.tourman.2022.104659.
T. Zhang, A. Koutsoumpis, J. K. Oostrom, D. Holtrop, S. Ghassemi, and R. E. de Vries, “Can Large Language Models Assess Personality from Asynchronous Video Interviews? A Comprehensive Evaluation of Validity, Reliability, Fairness, and Rating Patterns,” IEEE Trans Affect Comput, 2024, doi: 10.1109/TAFFC.2024.3374875.
Y. Bounab, M. Oussalah, N. Arhab, and S. Bekhouche, “Towards job screening and personality traits estimation from video transcriptions,” Expert Syst Appl, vol. 238, Mar. 2024, doi: 10.1016/j.eswa.2023.122016.
C. Qin et al., “Automatic Skill-Oriented Question Generation and Recommendation for Intelligent Job Interviews,” ACM Trans Inf Syst, vol. 42, no. 1, Aug. 2023, doi: 10.1145/3604552.
S. Chowdhury, P. Budhwar, and G. Wood, “Generative Artificial Intelligence in Business: Towards a Strategic Human Resource Management Framework,” British Journal of Management, 2024, doi: 10.1111/1467-8551.12824.
A. K. Upadhyay and K. Khandelwal, “Applying artificial intelligence: implications for recruitment,” Strategic HR Review, vol. 17, no. 5, pp. 255–258, Oct. 2018, doi: 10.1108/shr-07-2018-0051.
U. Leicht-Deobald et al., “The challenges of algorithm-based hr decision-making for personal integrity,” in Business and the Ethical Implications of Technology, Springer, 2022, pp. 71–86. doi: 10.1007/s10551-019-04204-w.
A. Ujlayan, S. Bhattacharya, and Sonakshi, “A Machine Learning-Based AI Framework to Optimize the Recruitment Screening Process,” International Journal of Global Business and Competitiveness, vol. 18, no. S1, pp. 38–53, Dec. 2023, doi: 10.1007/s42943-023-00086-y.
C. Fang, N. Markuzon, N. Patel, and J.-D. Rueda, “Patient-Reported Outcomes Natural Language Processing for Automated Classification of Qualitative Data From Interviews of Patients With Cancer,” vol. 25, no. 12, pp. 1995–2002, 2022, doi: 10.1016/j.
Y.-H. Chan and Y.-C. Fan, “A Recurrent BERT-based Model for Question Generation,” 2019.
S. Mahajan, N. Sonawane, R. Bagal, S. Kulkarni, A. Suryawanshi, and Y. Mhaisne, “GENERATIVE AI-BASED INTERVIEW SIMULATION AND PERFORMANCE ANALYSIS,” www.irjmets.com @International Research Journal of Modernization in Engineering, 2024, doi: 10.56726/IRJMETS56275.
A. Vaswani et al., “Attention Is All You Need,” Jun. 2017, [Online]. Available: http://arxiv.org/abs/1706.03762
J. Devlin, M.-W. Chang, K. Lee, K. T. Google, and A. I. Language, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” 2018. [Online]. Available: https://github.com/tensorflow/tensor2tensor
F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” Nov. 2020, [Online]. Available: http://arxiv.org/abs/2011.00677
T. B. Brown et al., “Language Models are Few-Shot Learners,” 2020. [Online]. Available: https://commoncrawl.org/the-data/
OpenAI, “GPT-4 Technical Report,” 2023.
J. Bai et al., “Qwen Technical Report,” Sep. 2023, [Online]. Available: http://arxiv.org/abs/2309.16609
S. Furniturewala et al., “‘Thinking’ Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models.”
S. Al Faraby, A. Romadhony, and Adiwijaya, “Analysis of LLMs for educational question classification and generation,” Computers and Education: Artificial Intelligence, vol. 7, Dec. 2024, doi: 10.1016/j.caeai.2024.100298.
M. Li et al., “EZInterviewer: To Improve Job Interview Performance with Mock Interview Generator,” in WSDM 2023 - Proceedings of the 16th ACM International Conference on Web Search and Data Mining, Association for Computing Machinery, Inc, Feb. 2023, pp. 1102–1110. doi: 10.1145/3539597.3570476.
H. L. Chung, Y. H. Chan, and Y. C. Fan, “Handover QG: Question Generation by Decoder Fusion and Reinforcement Learning,” IEEE/ACM Trans Audio Speech Lang Process, vol. 32, pp. 3644–3655, 2024, doi: 10.1109/TASLP.2024.3426292.
J. Thakkar, C. Thomas, and D. B. Jayagopi, “Automatic assessment of communication skill in real-world job interviews: A comparative study using deep learning and domain adaptation,” in ACM International Conference Proceeding Series, Association for Computing Machinery, Dec. 2023. doi: 10.1145/3627631.3627636.
K. Yadav, A. Seemendra, A. Singhania, S. Bora, P. Dubey, and V. Aggarwal, “Interviewing the Interviewer: AI-generated Insights to Help Conduct Candidate-centric Interviews,” in International Conference on Intelligent User Interfaces, Proceedings IUI, Association for Computing Machinery, Mar. 2023, pp. 723–736. doi: 10.1145/3581641.3584051.
C. Rigotti and E. Fosch-Villaronga, “Fairness, AI & recruitment,” Computer Law and Security Review, vol. 53, Jul. 2024, doi: 10.1016/j.clsr.2024.105966.
H. Sun et al., “Facilitating Multi-Role and Multi-Behavior Collaboration of Large Language Models for Online Job Seeking and Recruiting,” 2024. doi: XXXXXXX.XXXXXXX.
S. Lloyd, M. Beckman, D. Pearl, R. Passonneau, Z. Li, and Z. Wang, “Foundations for AI-Assisted Formative Assessment Feedback for Short-Answer Tasks in Large-Enrollment Classes,” in Bridging the Gap: Empowering and Educating Today’s Learners in Statistics. Proceedings of the Eleventh International Conference on Teaching Statistics, International Association for Statistical Education, Dec. 2022. doi: 10.52041/iase.icots11.T3C3.
N. A. Khayi, V. Rus, and L. Tamang, “Towards Improving Open Student Answer Assessment using Pretrained Transformers.”
DOI: https://doi.org/10.47738/jads.v7i1.970
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)