A Hybrid Method for Low-Resource Named Entity Recognition

Do Minh Duc, Quan Xuan Truong, Viet Tran Hong, Le Hoang Anh, Mac Thi Minh Tra, Nguyen Van Thuy, Le Hai Ha, Vinh Nguyen Van

Abstract


Named Entity Recognition (NER) is a critical component of Natural Language Processing with diverse applications in information extraction and conversational AI. However, NER in specific domains for low-resource languages faces challenges such as limited annotated data and heterogeneous label sets. This study addresses these issues by proposing a hybrid neurosymbolic framework that integrates rule-based processing with deep learning models for Vietnamese NER. The core idea involves a two-stage pipeline: first, a rule-based component reduces label complexity by grouping relational and special categories; second, pre-trained language models are fine-tuned for high-precision extraction. A post-processing module is then utilized to restore fine-grained labels, preserving expressiveness for application-level usability. To mitigate data scarcity, a scalable data augmentation strategy leveraging Large Language Models (LLMs) is introduced to expand the label set without full re-annotation—a significant novelty of this work. The effectiveness of this method was evaluated across five specific-domain datasets, including logistics, wildlife, and healthcare. Experimental results demonstrate substantial improvements over strong RoBERTa-based baselines. Specifically, the proposed system achieved F1 scores of 90% in Customer Service (up from 83%), 84% in GAM (up from 73%), 83% in AI Fluent (up from 80%), 94% in PhoNER_Covid19 (up from 91%), and 60% in Rare Wildlife (up from 36%). These findings confirm that the hybrid approach effectively captures the linguistic complexity of Vietnamese and contextual nuances in specialized domains, offering a robust contribution to low-resource NER research.

Keywords


Named Entity Recognition; Hybrid Model; Deep Learning; Rule-based System; Information Extraction

Full Text:

PDF

References


D. Nadeau and S. Sekine, “A survey of named entity recognition and classification,” Lingvisticae Investigationes, vol. 30, no. 1, pp. 3-26, 2007. https://doi.org/10.1075/li.30.1.03nad.

J. Li, A. Sun, J. Han, and C. Li, “A survey on deep learning for named entity recognition,” IEEE Trans. Knowl. Data Eng., vol. 34, no. 1, pp. 50-70, 2022. https://www.doi.org/10.1109/TKDE.2020.2981314.

J. Devlin, M-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. 2019 Conf. North Amer. Chapter Assoc. Comput. Linguistics, 2019. https://doi.org/10.18653/v1/N19-1423.

V. Yadav and S. Bethard, “A survey on recent advances in named entity recognition from deep learning models,” in Proc. 27th Int. Conf. Comput. Linguistics, 2019, pp. 2145-2158. https://doi.org/10.48550/arXiv.1910.11470.

I. Beltagy, K. Lo, and A. Cohan, “SciBERT: A Pretrained Language Model for Scientific Text,” in Proc. 2019 Conf. Empirical Methods Natural Language Processing, 2019. https://doi.org/10.18653/v1/D19-1371.

D.Q. Nguyen and A.T. Nguyen, “PhoBERT: Pre-trained language models for Vietnamese,” arXiv preprint arXiv:2003.00744, 2020. https://doi.org/10.48550/arXiv.2003.00744.

L.N. Chi, N.Y. Nguyen, and A.D. Trinh, “On the Vietnamese Name Entity Recognition: A Deep Learning Method Approach,” in RIVF Int. Conf. Computing Commun. Technol. (RIVF), 2019, pp. 1-5. https://www.doi.org/10.1109/RIVF48685.2020.9140754.

T.H. Truong, M. Dao, and D.Q. Nguyen, “COVID-19 Named Entity Recognition for Vietnamese,” 2021. https://doi.org/10.18653/v1/2021.naacl-main.173.

Y-T. Lu and Y. Huo, “Financial Named Entity Recognition: How Far Can LLM Go?,” in Proc. Joint Workshop 9th Financial Technology Natural Language Processing (FinNLP), 6th Financial Narrative Processing (FNP), and 1st Workshop Large Language Models Finance Legal (LLMFinLegal), 2025. https://aclanthology.org/2025.finnlp-1.15/.

E. Tjong Kim Sang and F. De Meulder, “Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition,” in Proc. Seventh Conf. Natural Language Learning, 2003. https://doi.org/10.48550/arXiv.cs/0306050.

R. Grishman and B.M. Sundheim, “Message Understanding Conference-6: A Brief History,” in Proc. 16th Int. Conf. Comput. Linguistics, 1996. https://www.doi.org/10.3115/992628.992709.

J. Guo, G. Xu, X. Cheng, and H. Li, “Named entity recognition in query,” in Proc. 32nd Int. ACM SIGIR Conf. Research Dev. Information Retrieval, 2009. https://www.doi.org/10.1145/1571941.1571989.

D.I. Moldovan, M. Pasca, S.M. Harabagiu, and M. Surdeanu, “Performance Issues and Error Analysis in an Open-Domain Question Answering System,” ACM Trans. Inf. Syst., vol. 21, pp. 133-154, 2002. https://doi.org/10.3115/1073083.1073091.

B. Babych and A. Hartley, “Improving machine translation quality with automatic named entity recognition,” in Proc. EAMT-ISTAS Workshop, 2003, pp. 1-9. https://www.doi.org/10.3115/1609822.1609823.

H. Ji and R. Grishman, “Knowledge Base Population: Successful Approaches and Challenges,” in Proc. 49th Ann. Meet. Assoc. Comput. Linguistics: Human Lang. Technol., Portland, OR, USA, 2011, pp. 1148–1158. https://aclanthology.org/P11-1115/.

G. Lample, M. Ballesteros, S. Subramanian, K. Kawakami, C. Dyer, “Neural architectures for named entity recognition,” in Proc. 2016 Conf. North Amer. Chapter Assoc. Comput. Linguistics: Human Lang. Technol., 2016, pp. 260-270. https://doi.org/10.18653/v1/N16-1030.

G.K.O. Crichton, S. Pyysalo, B. Chiu, and A. Korhonen, “A neural network multi-task learning approach to biomedical named entity recognition,” BMC Bioinformatics, vol. 18, 2017. https://doi.org/10.1186/s12859-017-1776-8.

Y. Wang, L. Wang, M. Rastegar-Mojarad, H. Liu, “Cross-type biomedical named entity recognition with deep multi-task learning,” Bioinformatics, vol. 35, no. 10, pp. 1745–1752, 2019. https://doi.org/10.1093/bioinformatics/bty869.

V. Pais, M. Mitrofan, C.L. Gasan, V. Coneschi, A. Ianov, “Named Entity Recognition in the Romanian Legal Domain,” in Proc. Natural Legal Language Processing Workshop 2021, 2021, pp. 9–18. https://doi.org/10.18653/v1/2021.nllp-1.2.

V. Naik, P. Patel, R. Kannan, “Legal Entity Extraction: An Experimental Study of NER Approach for Legal Documents,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 3, pp. 775–781, 2023. https://www.doi.org/10.14569/IJACSA.2023.0140389.

A. Shah, A. Gullapalli, R. Vithani, M. Galarnyk, S. Chava, “FiNER-ORD: Financial Named Entity Recognition Open Research Dataset,” 2023. https://doi.org/10.48550/arXiv.2302.11157.

P.Q.N. Minh, “A Feature-Rich Vietnamese Named-Entity Recognition Model,” 2018. https://doi.org/10.48550/arXiv.1803.04375.

H. Wu and K. Tu, “Layer-Condensed KV Cache for Efficient Inference of Large Language Models,” in Proc. 62nd Ann. Meet. Assoc. Comput. Linguistics, 2024. https://doi.org/10.18653/v1/2024.acl-long.602.

J. Dai, Z. Huang, H. Jiang, C. Chen, D. Cai, et al., “CORM: Cache Optimization with Recent Message for Large Language Model Inference,” 2024. https://doi.org/10.48550/arXiv.2404.15949.

T. Dao, D.Y. Fu, S. Ermon, A. Rudra, C. Ré, “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” arXiv, 2022. https://doi.org/10.48550/arXiv.2205.14135.

Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” 2019. https://doi.org/10.48550/arXiv.1907.11692.

D.P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” CoRR, 2014. https://doi.org/10.48550/arXiv.1412.6980.




DOI: https://doi.org/10.47738/jads.v7i2.1161

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0