A Hybrid Fuzzy-LLM Framework for Difficulty Estimation of Math Word Problems: A Data-Driven Human-in-the-Loop Study

Shilpa Kadam, Jabez Christopher, PTV Praveen Kumar, Dipak Kumar Satpathi

Abstract


Assessing the difficulty levels of Math Word Problems (MWPs) is essential for adaptive learning, yet most existing MWP datasets lack standardized difficulty annotations. This paper proposes a decision framework that integrates a 2-tuple Fuzzy Linguistic Decision Model (FLDM) with Large Language Models (LLMs) for automated difficulty estimation. A corpus of over 2,000 MWPs was compiled, of which 200 were annotated by seven instructors and an additional 454 were validated by ten experts. Consensus stability improved markedly (Fleiss’ κ = 0.14 → Cohen’s κ = 0.32), reflecting stronger alignment between expert judgments and the proposed fuzzy 2-tuple aggregation. Sixteen LLM configurations were evaluated, including GPT-3.5, GPT-4o-Mini, Gemini Flash, and LLaMA-3.2 under Zero-Shot, Five-Shot, and RAG settings. GPT-3.5 Zero-Shot achieved the best performance (Precision=0.65, Recall=0.63, F1=0.63), outperforming GPT-4o-Mini and Gemini variants. The validated dataset and linguistic ground truth were integrated into a web-based annotation system (themathbits.com), demonstrating scalability for real-world deployment. The results show that combining human linguistic judgments with fuzzy modeling and LLM inference improves reliability of MWP difficulty estimation, providing a foundation for future adaptive learning platforms. 


Keywords


math word problems; difficulty estimation; fuzzy linguistic decision model; large language models; educational data; expert annotation; adaptive learning

Full Text:

PDF

References


R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman and H. Hajishirzi, “{MAWPS}: A Math Word Problem Repository,” in Proceedings of the 2016 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego, California, 2016.

G. Daroczy, M. Wolska, W. D. Meurers and H. C. Nuerk, “Word problems: A review of linguistic and numerical factors contributing to their difficulty,” Frontiers in Psychology, vol. 6, 2015.

L. Verschaffel, S. Schukajlow, J. Star and W. V. Dooren, “Word problems in mathematics education: a survey,” ZDM - Mathematics Education, vol. 52, 2020.

Z. Ding, X. Wang, Y. Wu, G. Cao and L. Chen, “Tagging knowledge concepts for math problems based on multi-label text classification,” Expert Systems with Applications, vol. 267, 2025.

S. Kadam, P. K. Srungaram, S. D. Y, M. S.S.S.R, P. Praveen, S. Pappu and D. K. Satpathi, “Analysis of Linguistics and Math Features for Classification of Math Word Problems: Insights and Future Direction,” International Journal of Management and Applied Science (IJMAS)-IJMAS, 2023.

E. M.-R. a. E. Hernández-Pereira, D. Alonso-Ríos, J. Bobes-Bascarán and Á. Fernández-Leal, “Human-in-the-loop machine learning: a state of the art,” Artificial Intelligence Review, vol. 56, 2023.

Y. Zhang, Y. Luo, Y. Yuan and A. C.-C. Yao, “Autonomous Data Selection with Language Models for Mathematical Texts,” 2024.

X. Y. a. X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su and W. Chen, “MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning,” arXiv, 2023.

L. Y. a. W. J. a. H. S. a. J. Y. a. Z. Liu, Y. Zhang, J. T. Kwok, Z. L. a. A. Weller and W. Liu, “MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models,” arXiv, 2023.

S. Mishra, M. Finlayson, P. Lu, L. Tang, S. Welleck, C. Baral, T. Rajpurohit, O. Tafjord, A. Sabharwal, P. Clark and A. Kalyan, “LĪLA: A Unified Benchmark for Mathematical Reasoning,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, 2022.

K. C. Schulman, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse and J. Schulman, “Training Verifiers to Solve Math Word Problems,” arXiv, 2021.

A. Patel, S. Bhattamishra and N. Goyal, “Are NLP Models really able to Solve Simple Math Word Problems?,” in NAACL-HLT 2021 - 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 2021.

R. R. M and L. Martínez, “An analysis of symbolic linguistic computing models in decision making,” International Journal of General Systems, vol. 42, no. 1, pp. 121-136, 2013.

K. Nate, A. Yoav, Z. Luke and BarzilayRegina, “Learning to automatically solve algebra word problems,” in 52nd Annual Meeting of the Association for Computational Linguistics, ACL 2014 - Proceedings of the Conference, 2014.

S.-y. Miao, C.-C. Liang and K.-Y. Su, “A Diverse Corpus for Evaluating and Developing {E}nglish Math Word Problem Solvers,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020.

D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. T. a. D. Song and J. Steinhardt, “Measuring Mathematical Problem Solving With the MATH Dataset,” NeurIPS 2021, 2021.

F. Herrera and L. Martínez, “An approach for combining linguistic and numerical information based on the 2-tuple fuzzy linguistic representation model in decision-making,” International Journal of Uncertainty, Fuzziness and Knowldege-Based Systems, vol. 8, pp. 539-562, 2000.

R. A. Carrasco, M. F. Blasco and E. Herrera-Viedma, “A 2-tuple fuzzy linguistic RFM model and its implementation,” Procedia Computer Science, vol. 55, pp. 1340-1347, 2015.

L. Zadeh, “The concept of a linguistic variable and its application to approximate reasoning—I,” Information Sciences, vol. 8, 1975.

F. Herrera and E. Herrera-Viedma, “Choice functions and mechanisms for linguistic preference relations,” European Journal of Operational Research, vol. 120, pp. 141-161, 2000.

F. Soygazi and D. Oguz, “An Analysis of Large Language Models and LangChain in Mathematics Education,” Association for Computing Machinery, 2024.

F. D. a. M. Herrmann, “Using large language models to support pre-service teachers mathematical reasoning—an exploratory study on ChatGPT as an instrument for creating mathematical proofs in geometry.,” Front. Artif. Intell, 2024.

Z. Levonian, C. Li, W. Zhu, A. Gade, O. Henkel, M.-E. Postle and W. Xing, “Retrieval-augmented Generation to Improve Math Question-Answering: Trade-offs Between Groundedness and Human Preference,” 2023.

K. Zaporojets, G. Bekoulis, J. Deleu, T. Demeester and C. Develder, “Solving arithmetic word problems by scoring equations with recursive neural networks,” Expert Systems with Applications, vol. 174, 2021.

R. Meissner, A. Pögelt, K. Ihsberner, M. Grüttmüller, S. Tornack, A. Thor, N. Pengel, H.-W. Wollersheim and W. Hardt, “LLM-generated competence-based e-assessment items for higher education mathematics: methodology and evaluation,” Frontiers in Education, vol. 9, 2024.

F. Herrera and E. Herrera-Viedma, “Linguistic decision analysis: steps for solving decision problems under linguistic information,” Fuzzy Sets and Systems, vol. 15, pp. 67-82, 2000.

F. Herrera and L. Martínez, “A Model Based on Linguistic 2-Tuples for Dealing with Multigranular Hierarchical Linguistic Contexts in Multi-Expert Decision-Making,” CYBERNETICS, vol. 31, no. 2, 2001.

A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y. Choi and H. Hajishirzi, “MathQA: Towards interpretable math word problem solving with operation-based formalisms,” in NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 2019.

T. Malhotra and A. Gupta, “A systematic review of developments in the 2-tuple linguistic model and its applications in decision analysis,” Soft Computing, vol. 27, no. 3, pp. 1871-1905, 2023.

C.-K. Law, “Using fuzzy numbers in educational grading system,” Fuzzy Sets and Systems, vol. 83, no. 3, pp. 311-323, 1996.

H.-M. Lee, “Group decision making using fuzzy sets theory for evaluating the rate of aggregative risk in software development,” Fuzzy Sets and Systems, vol. 80, no. 3, pp. 261-271, 1996.

F. Herrera, E. Herrera-Viedma, S. Alonso and F. Chiclana, “Computing with words in decision making: foundations, trends and prospects,” Fuzzy Optimization and Decision Making, 2009.

L. Martínez, R. M. Rodriguez and F. Herrera, The 2-tuple Linguistic Model: Computing with Words in Decision Making, Cham: Springer International Publishing, 2015, pp. 131--143.

L. Martínez, “Computing with words in linguistic decision making: Analysis of linguistic computing models,” Proceedings of 2010 IEEE International Conference on Intelligent Systems and Knowledge Engineering, ISKE 2010, pp. 5-8, 2010.

F. Herrera, E. Herrera-Viedma and J. Verdegay, “A rational consensus model in group decision making using linguistic assessments,” Fuzzy Sets and Systems, vol. 88, pp. 31-49, 1997.

L. M. a. F. Herrera, “2-Tuple linguistic representation model,Computing with words,Fuzzy linguistic approach,Linguistic variable,” Information Sciences, vol. 207, pp. 1-18, 11 2012.

F. Herrera and L. Martínez, “Computing with words (CW), information fusion, linguistic modeling, linguistic variables,” IEEE Transactions on Fuzzy Systems, vol. 8, no. 6, 2000.

Z. Xu and H. Wang, “On the syntax and semantics of virtual linguistic terms for information fusion in decision making,” Information Fusion, vol. 34, pp. 43-48, 3 2017.

D. DuBois and H. Prade, Fuzzy Sets and Systems: Theory and Applications, Academic Press, Inc, 1997.

Y. R. R., “An approach to ordinal decision making,” Int. J. Approx. Reason., vol. 12, pp. 237-261, 1995.

A. Scarlatos and A. Lan, “Tree-Based Representation and Generation of Natural and Mathematical Language,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers, 2023.

J. Jaeho and L. Seongyong, “Large language models in education: A focus on the complementary relationship between human teachers and ChatGPT,” Education and Information Technologies, 2023.

J. Ahn, R. Verma, R. Lou, D. Liu, R. Zhang and W. Yin, “Large Language Models for Mathematical Reasoning: Progresses and Challenges,” 2024.

H. M. Javad, H. Hannaneh and E. O. a. K. Nate, “Learning to solve arithmetic word problems with verb categorization,” in EMNLP 2014 - 2014 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference}, 2014.

Y. Wang, X. Liu and S. Shi, “Deep neural solver for math word problems,” in EMNLP 2017 - Conference on Empirical Methods in Natural Language Processing, Proceedings, 2017.

L. W. :. Y. D. :. D. C. :. B. Phil, “Program induction by rationale generation: Learning to solve and explain algebraic word problems,” in ACL 2017 - 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers), 2017.

X. Sun, X. Li, J. Li, F. Wu, S. Guo, T. Zhang and G. Wang, “Text Classification via Large Language Models,” 2023.




DOI: https://doi.org/10.47738/jads.v7i2.1187

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0