A Study to Detect Multi-word Expression from Text Using Deep Learning Models

Wong Jun Meng, Tan Yu Jie, Lim Tong Ming

Abstract


Detecting Multi-word Expressions (MWEs) is a crucial task in Natural Language Processing (NLP) for applications in machine translation, sentiment analysis, and information retrieval. This study evaluates the performance of several deep learning models on MWE detection using two samples of varying sizes from the major consumer electronic product retailer corpus. The sample is limited to 10,000 and 15,000 rows, with each row contains 15-20 English words. Preprocessing steps include removing special symbols and emojis, converting text to lowercase, and applying the spaCy NLP library for tokenization and part-of-speech (POS) tagging. Syntactic rules are then used to identify MWEs such as verb-noun combinations and phrasal verbs, with BIO tags (B-MWE, I-MWE, O) to mark MWE boundaries. We investigated transformer-based models such as BERT, BERT-CRF, LSTM-CRF and RoBERTa-CRF using a sample of 10,000 rows; BERT, BERT-BiLSTM, BiLSTM-GloVe, and BiLSTM-GloVe-BiGRU uses a sample of 15,000. Results demonstrated that the transformer-based model, RoBERTa-CRF, excels on the smaller sample which achieves the best performance by leveraging the contextual embeddings and sequential dependency modeling. On a larger sample, the BERT-BiLSTM model emerged as the most effective model, showcasing the advantage of combining dynamic embeddings with sequential learning. In contrast, models utilizing static embeddings, such as GloVe, displayed moderate performance, highlighting their limitations in capturing contextual nuances. Comparative analysis across both samples reveals that transformer-based models like RoBERTa-CRF performed optimally on the smaller dataset, whereas hybrid models integrating with sequential architectures like BERT-BiLSTM demonstrated superior performance as dataset size increased. These findings highlight the importance of model selection based on dataset scale to optimize MWE detection. This study underscores the importance of integrating contextual and sequential deep learning techniques to improve MWE detection and provides a basis for developing more robust and scalable systems for diverse linguistic tasks.


Keywords


Multiword Expressions; Natural Language Processing; Deep Learning; Transformer Models; Recurrent Neural Networks; Contextual Embeddings; Sequence Labeling

Full Text:

PDF

References


References

Constant, M., Eryiğit, G., Monti, J., van der Plas, L., Ramisch, C., Rosner, M., & Todirascu, A. (2017). Multiword Expression Processing: A Survey. Computational Linguistics, 43(4), 837–892. https://doi.org/10.1162/coli_a_00302

Jisha, T., & Thomas, M. (2024). Identification of Multiword Expressions: A Literature Study. Retrieved December 15, 2024, from https://marymathacollege.ac.in/data/downloads/2019-01-21-2-25-58_identification-of-multi-word-expressions-jisha.pdf

Gharbieh, W., Bhavsar, V., & Cook, P. (2017). Deep Learning Models For Multiword Expression Identification (pp. 54–64). Association for Computational Linguistics. https://aclanthology.org/S17-1006.pdf

Villena-Román, J., Collada-Pérez, S., Lana-Serrano, S., & González, C. (2011). Hybrid Approach Combining Machine Learning and a Rule-Based Expert System for Text Categorization. Proceedings of the Twenty-Fourth International Florida Artificial Intelligence Research Society Conference, May 18-20, 2011, Palm Beach, Florida, USA. https://www.researchgate.net/publication/221438964_Hybrid_Approach_Combining_Machine_Learning_and_a_Rule-Based_Expert_System_for_Text_Categorization

Islam, S., Hanae Elmekki, Elsebai, A., Bentahar, J., Nagat Drawel, Gaith Rjoub, & Witold Pedrycz. (2023). A comprehensive survey on applications of transformers for deep learning tasks. Expert Systems with Applications, 122666–122666. https://doi.org/10.1016/j.eswa.2023.122666

Jogin, M., Mohana, Madhulika, M. S., Divya, G. D., Meghana, R. K., & Apoorva, S. (2018). Feature Extraction using Convolution Neural Networks (CNN) and Deep Learning. IEEE Xplore. https://doi.org/10.1109/RTEICT42901.2018.9012507

Schmidhuber, J. (2015). Deep learning in neural networks: An overview. Neural Networks, 61(61), 85–117. https://doi.org/10.1016/j.neunet.2014.09.003

Sainath, T. N., Vinyals, O., Senior, A., & Sak, H. (2015, April 1). Convolutional, Long Short-Term Memory, fully connected Deep Neural Networks. IEEE Xplore. https://doi.org/10.1109/ICASSP.2015.7178838

Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2020). Generative adversarial networks. Communications of the ACM, 63(11), 139–144. https://doi.org/10.1145/3422622

Chen, S., & Guo, W. (2023). Auto-Encoders in Deep Learning—A Review with New Perspectives. Mathematics, 11(8), 1777. https://doi.org/10.3390/math11081777

Damith Premasiri, & Ranasinghe, T. (2022, August 16). BERT(s) to Detect Multiword Expressions. https://doi.org/10.48550/arXiv.2208.07832

Savary, A., Khelil, C., Ramisch, C., Giouli, V., Barbu, V., Racai, M., Academy, R., Hadj, N., Krstev, C., Liebeskind, C., Xu, H., Jiang, M., Stymne, S., Pickard, T., Guillaume, B., Bhatia, A., Butler, A., Candito, M., Gantar, A., & Jožef, S. (2023). PARSEME Corpus Release 1.3 (pp. 24–35). https://aclanthology.org/2023.mwe-1.6.pdf

dimsum16. (2015, December 28). GitHub - dimsum16/dimsum-data: Data for the DiMSUM shared task at SEMEVAL 2016. GitHub. https://github.com/dimsum16/dimsum-data

Ramisch, C., Walsh, A., Blanchard, T., & Taslimipoor, S. (2023). A Survey of MWE Identification Experiments: The Devil is in the Details (pp. 106–120). https://aclanthology.org/2023.mwe-1.15.pdf

Tuora, R. (2021). Dependency Trees in Automatic Inflection of Multiword Expressions in Polish. CLARIN Annual Conference 2021. https://office.clarin.eu/v/CE-2021-1923-CLARIN2021_ConferenceProceedings.pdf

Erden, B. (2019). Identification of verbal multiword expressions using deep learning architectures and representation learning methods (Doctoral dissertation, Bogaziçi University). https://www.cmpe.boun.edu.tr/~gungort/theses/Identification%20of%20Verbal%20Multiword%20Expressions%20using%20Deep%20Learning%20Architectures%20and%20Representation%20Learning%20Methods.pdf

Author name / JADS 00 (2019) 000–000

Taslimipoor, S., Rohanian, O., & Ha, L. (n.d.a). Cross-lingual Transfer Learning and Multitask Learning for Capturing Multiword Expressions. Retrieved December 19, 2024, from

https://wlv.openrepository.com/bitstream/handle/2436/623211/Taslimipoor_et_al_Cross-lingual_transfer_2019.pdf

Aziz, A., Hossain, M. A., Chy, A. N., Ullah, M. Z., & Aono, M. (2023). Leveraging contextual representations with BiLSTM-based regressor for lexical complexity prediction. Natural Language Processing Journal, 5, 100039. https://doi.org/10.1016/j.nlp.2023.100039

Garg, K. D., Shekhar, S., Kumar, A., Goyal, V., Sharma, B., Chengoden, R., & Srivastava, G. (2022). Framework for handling rare word problems in Neural Machine Translation System using Multi-Word Expressions. Applied Sciences, 12(21), 11038. https://doi.org/10.3390/app122111038

Vacareanu, R., Valenzuela-Escaŕcega, M., Sharp, R., & Surdeanu, M. (2020). An Unsupervised Method for Learning Representations of Multi-word Expressions for Semantic Classification (pp. 3346–3356). https://aclanthology.org/2020.coling-main.297.pdf

Mahtab Sarlak, Yalda Yarandi, & Mehrnoush Shamsfard. (2023). Predicting Compositionality of Verbal Multiword Expressions in Persian. 14–23. https://doi.org/10.18653/v1/2023.mwe-1.5

Taslimipoor, S., Rohanian, O., & Ha, L. (n.d.b). Cross-lingual Transfer Learning and Multitask Learning for Capturing Multiword Expressions. Retrieved December 19, 2024. https://wlv.openrepository.com/bitstream/handle/2436/623211/Taslimipoor_et_al_Cross-lingual_transfer_2019.pdf

Premasiri, D., Haddad, A. H., Ranasinghe, T., & Mitkov, R. (2022). Transformer-based detection of multiword expressions in flower and plant names. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2209.08016

Piasecki, M., & Kanclerz, K. (2022). Non-Contextual vs Contextual Word Embeddings in Multiword Expressions Detection. In Lecture notes in computer science (pp. 193–206). https://drive.google.com/file/d/1FkqgMYWJtjS9In8cTwDOd4htOFs8fEsp/view

Berend, G. (2018). l1 Regularization of Word Embeddings for Multi-Word Expression Identification. Acta Cybernetica, 23(3), 801–813. https://doi.org/10.14232/actacyb.23.3.2018.5




DOI: https://doi.org/10.47738/jads.v6i3.716

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN : 2723-6471 (Online)
Publisher : Bright Publisher
Website : http://bright-journal.org/JADS
Email : taqwa@amikompurwokerto.ac.id (principal contact)
    support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0