PRAKE: A Modified RAKE Model for Keyword Extraction in Accreditation Assessment Descriptions

Helena Nurramdhani Irmanda, Sri Hartati, Sri Mulyana

Abstract


Study program accreditation requires aligning assessment criteria with the Self-Evaluation Sheet (LED), which is usually written as a lengthy and complex narrative. Finding relevant information requires a method that can automatically extract keywords from assessment descriptions as representations of the criteria. Keyword extraction can be applied through the Rapid Automatic Keyword Extraction (RAKE) method, a simple technique that works without labeled data. However, standard RAKE uses stopwords as delimiters to segment candidate phrases, making it less effective for complex sentences such as those found in accreditation assessment descriptions. Because a single sentence may contain several ideas, the extraction process should handle phrases carefully through splitting, merging, or extension according to their structure and meaning. To address this limitation, this study introduces PRAKE (Phrase-Refined RAKE), a modified RAKE algorithm that enhances candidate phrase extraction. Modifications are carried out at the Candidate Phrase Extraction stage through three techniques, including Phrase Completion to complete short phrases afterwards with the prefix of the previous phrase, Phrase Restructuring to rearrange phrases through merging or separation based on structure and meaning, and Semantic Phrase Composition to form new phrases from different elements that are semantically interrelated. Additionally, a domain term weighting based on term frequency is integrated into the scoring calculation to strengthen the relevance of terms to the accreditation context. The model achieved a precision of 0.90, recall of 0.83, and F1-score of 0.85, representing the average performance across all 101 assessment descriptions evaluated in this study. The results demonstrate that PRAKE adapts better to accreditation terminology and improves keyword relevance and extraction efficiency. These findings indicate that PRAKE provides a foundation for automated evaluation and can be extended for cross-domain document analysis.

Keywords


Keyword Extraction; Rake; Study Program Accreditation; Natural Language Processing; Text Processing

Full Text:

PDF

References


L. Sudianto and P. Simon, “Application of Monitoring Database for Accreditation Instrument UKI PAULUS,” IOP Conf Ser Mater Sci Eng, vol. 846, no. 1, p. 12027, May 2020, doi: 10.1088/1757-899X/846/1/012027.

LAMINFOKOM, “Academic Paper for Study Program Accreditation,” 2021. Accessed: Feb. 28, 2024. [Online]. Available: https://laminfokom.or.id/official/img/instrumen/instrumen_2372Lampiran%201%20PerBAN-PT%2015%202021%20Instrumen%20APS%20Sarjana%20Infokom.pdf

A. Mulyanto, S. Hartati, and R. Wardoyo, “An Integrated Model of Natural Language Processing Technique and Case-Based Reasoning for Supporting Study Program Accreditation,” ICIC Express Letters, vol. 18, no. 7, pp. 749–757, 2024, doi: 10.24507/icicel.18.07.749.

S. Mulyana, S. Hartati, R. Wardoyo, and Subandi, “A Processing Model Using Natural Language Processing (NLP) For Narrative Text Of Medical Record For Producing Symptoms Of Mental Disorders,” in 2019 Fourth International Conference on Informatics and Computing (ICIC), 2019, pp. 1–6. doi: 10.1109/ICIC47613.2019.8985862.

Z. H. Amur, Y. K. Hooi, G. M. Soomro, H. Bhanbhro, S. Karyem, and N. Sohu, “Unlocking the Potential of Keyword Extraction: The Need for Access to High-Quality Datasets,” Applied Sciences, vol. 13, no. 12, p. 7228, Jun. 2023, doi: 10.3390/app13127228.

S. Rose, D. Engel, N. Cramer, and W. Cowley, “Automatic keyword extraction from individual documents,” Text mining: applications and theory, pp. 1–20, 2010, doi: 10.1002/9780470689646.ch1.

M. Nadim, D. Akopian, and A. Matamoros, “A Comparative Assessment of Unsupervised Keyword Extraction Tools,” J Inf Sci, vol. 49, no. 4, pp. 573–588, 2023, doi: 10.1109/ACCESS.2023.3344032.

V. Singh and B. K. Bolla, “Hybrid Approach To Unsupervised Keyphrase Extraction,” Procedia Comput Sci, vol. 235, pp. 1498–1511, 2024, doi: 10.1016/j.procs.2024.04.141.

S. G. Jindal and A. Kaur, “Automatic Keyword and Sentence-Based Text Summarization for Software Bug Reports,” IEEE Access, vol. 8, pp. 65352–65370, 2020, doi: 10.1109/ACCESS.2020.2985222.

K. Rinartha and L. G. S. Kartika, “Rapid automatic keyword extraction and word frequency in scientific article keywords extraction,” in 2021 3rd International conference on cybernetics and intelligent system (ICORIS), 2021, pp. 1–4. doi: 10.1109/ICORIS52787.2021.9649458.

G. Muppala and T. Devi, “Accurate Recasting of Giant Text into Charts Using Rapid Automatic Keyword Extraction Algorithm in Comparison with Bag of Words Algorithm,” in 2023 6th International Conference on Contemporary Computing and Informatics (IC3I), 2023, pp. 2548–2552. doi: 10.1109/IC3I59117.2023.10397804.

E. Papagiannopoulou and G. Tsoumakas, “A review of keyphrase extraction,” Wiley Interdiscip Rev Data Min Knowl Discov, vol. 10, no. 2, p. e1339, 2020, doi: 10.1002/widm.1339.

R. Campos, V. Mangaravite, A. Pasquali, A. Jorge, C. Nunes, and A. Jatowt, “YAKE! Keyword extraction from single documents using multiple local features,” Inf Sci (N Y), vol. 509, pp. 257–289, 2020, doi: 10.1016/j.ins.2019.09.013.

R. Mihalcea and P. Tarau, “Textrank: Bringing order into text,” in Proceedings of the 2004 conference on empirical methods in natural language processing, 2004, pp. 404–411.

N. Firoozeh, A. Nazarenko, F. Alizon, and B. Daille, “Keyword extraction: Issues and methods,” Nat Lang Eng, vol. 26, no. 3, pp. 259–291, 2020, doi: 10.1017/S1351324919000457.

L. Ajallouda, F. Z. Fagroud, A. Zellou, and E. habib Benlahmar, “Automatic keyphrases extraction: an overview of deep learning approaches,” Bulletin of Electrical Engineering and Informatics, vol. 12, no. 1, pp. 303–313, 2023, doi: 10.11591/eei.v12i1.4130.

Z. Alami Merrouni, B. Frikh, and B. Ouhbi, “Automatic keyphrase extraction: a survey and trends,” J Intell Inf Syst, vol. 54, no. 2, pp. 391–424, 2020, doi: 10.1007/s10844-019-00558-9.

B. Xie et al., “From statistical methods to deep learning, automatic keyphrase prediction: A survey,” Inf Process Manag, vol. 60, no. 4, p. 103382, 2023, doi: 10.1016/j.ipm.2023.103382.

H. Shin, H. J. Lee, and S. Cho, “General-use unsupervised keyword extraction model for keyword analysis,” Expert Syst Appl, vol. 233, p. 120889, 2023, doi: 10.1016/j.eswa.2023.120889.

M. Zhang, X. Li, S. Yue, and L. Yang, “An empirical study of TextRank for keyword extraction,” IEEE access, vol. 8, pp. 178849–178858, 2020, doi: 10.1109/ACCESS.2020.3027567.

C. Florescu and C. Caragea, “Positionrank: An unsupervised approach to keyphrase extraction from scholarly documents,” in Proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: long papers), 2017, pp. 1105–1115. doi: 10.18653/v1/P17-1102.

O. Alqaryouti, H. Khwileh, T. Farouk, A. Nabhan, and K. Shaalan, “Graph-based keyword extraction,” Intelligent Natural Language Processing: Trends and Applications, pp. 159–172, 2018, doi: 10.1007/978-3-319-67056-0_9.

J. Sawicki, M. Ganzha, and M. Paprzycki, “The state of the art of natural language processing—a systematic automated review of NLP literature using NLP techniques,” Data Intell, vol. 5, no. 3, pp. 707–749, 2023, doi: 10.1162/dint_a_00213.

K. Kanclerz et al., “What if ground truth is subjective? personalized deep neural hate speech detection,” in Proceedings of the 1st Workshop on Perspectivist Approaches to NLP@ LREC2022, 2022, pp. 37–45.

C. P. Chai, “Comparison of text preprocessing methods,” Nat Lang Eng, vol. 29, no. 3, pp. 509–553, 2023, doi: 10.1017/S1351324922000213.

Y. HaCohen-Kerner, D. Miller, and Y. Yigal, “The influence of preprocessing on text classification using a bag-of-words representation,” PLoS One, vol. 15, no. 5, p. e0232525, 2020, doi: 10.1371/journal.pone.0232525.

P. M. Rahate and M. Chandak, “An experimental technique on text normalization and its role in speech synthesis,” International Journal of Innovative Technology and Exploring Engineering (IJITEE), vol. 8, no. 8S3, pp. 545–548, 2019.

D. J. Ladani and N. P. Desai, “Stopword Identification and Removal Techniques on TC and IR applications: A Survey,” in 2020 6th International Conference on Advanced Computing and Communication Systems (ICACCS), Mar. 2020, pp. 466–472. doi: 10.1109/ICACCS48705.2020.9074166.

A. D. Latief, T. Sampurno, A. O. Arisha, and others, “Next sentence prediction: the impact of preprocessing techniques in deep learning,” in 2023 International Conference on Computer, Control, Informatics and its Applications (IC3INA), 2023, pp. 274–278. doi: 10.1109/IC3INA60834.2023.10285805.

P. Qi, Y. Zhang, Y. Zhang, J. Bolton, and C. D. Manning, “Stanza: A Python natural language processing toolkit for many human languages,” arXiv preprint arXiv:2003.07082, 2020, doi: 10.48550/arXiv.2003.07082.

J. Petrus, Ermatita, Sukemi, and Erwin, “An adaptable sentence segmentation based on Indonesian rules,” IAES International Journal of Artificial Intelligence, vol. 12, no. 3, pp. 1491 – 1499, 2023, doi: 10.11591/ijai.v12.i3.pp1491-1499.

M. A. Rosid, A. S. Fitrani, I. R. I. Astutik, N. I. Mulloh, and H. A. Gozali, “Improving text preprocessing for student complaint document classification using sastrawi,” in IOP Conference Series: Materials Science and Engineering, 2020, p. 12017. doi: 10.1088/1757-899X/874/1/012017.

S. Senbel, “Fast and Memory-Efficient TFIDF Calculation for Text Analysis of Large Datasets,” Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 12798 LNAI, pp. 557 – 563, 2021, doi: 10.1007/978-3-030-79457-6_47.

L. Afuan, N. Hidayat, N. Nofiyati, and M. F. As’ ad, “Sentiment Analysis of the Kampus Merdeka Program on Twitter Using Support Vector Machine and a Feature Extraction Comparison: TF-IDF vs. FastText,” Journal of Applied Data Sciences, vol. 5, no. 4, pp. 1738–1753, 2024, doi: 10.47738/jads.v5i4.436.

C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge University Press, 2008.

H. N. Irmanda and S. Hartati, “Sentiment Analysis of Cyberbullying Using Machine Learning,” in 2024 International Conference on Informatics, Multimedia, Cyber and Information System (ICIMCIS), 2024, pp. 594–600. doi: 10.1109/ICIMCIS63449.2024.10957620.

J. Shen et al., “Utilizing Natural Language Processing for Efficient Text Analysis in the Era of Social Media,” in Proceedings of 2024 IEEE 25th China Conference on System Simulation Technology and its Application, CCSSTA 2024, 2024, pp. 320 – 325. doi: 10.1109/CCSSTA62096.2024.10691866.

S. Akter, F. M. J. M. Shamrat, S. Chakraborty, A. Karim, and S. Azam, “COVID-19 Detection Using Deep Learning Algorithm on Chest X-ray Images,” Biology (Basel), vol. 10, no. 11, 2021, doi: 10.3390/biology10111174.

P. Mishra, C. M. Pandey, U. Singh, A. Gupta, C. Sahu, and A. Keshri, “Descriptive statistics and normality tests for statistical data,” Ann Card Anaesth, vol. 22, no. 1, pp. 67–72, 2019, doi: 10.4103/aca.ACA_157_18.




DOI: https://doi.org/10.47738/jads.v7i2.1057

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0