Comparison of Multilingual Model Sensitivity for Political Fact Verification with Integrated Multi-Evidence

Nova Agustina, Kusrini Kusrini, Ema Utami, Tonny Hidayat

Abstract


Political news is frequently targeted by the dissemination of fake news on social media, which can influence public opinion and undermine trust in democratic processes. The main challenge in addressing this issue lies in the limited sensitivity of cross-lingual fact verification models in capturing semantic relationships between claims and evidence in long-text, multi-evidence settings. Existing approaches often struggle to assess the relevance and quality of evidence, resulting in suboptimal verification performance. This study compares three multilingual Large Language Models (LLMs), namely mBERT, XLM-R, and LaBSE, for political fact verification using an integrated multi-evidence approach. Experiments are conducted on the PolitiFact dataset, with performance evaluated using sensitivity, accuracy, precision, and F1-score metrics.The results indicate that mBERT achieves the highest overall sensitivity at 89.44%, followed by LaBSE at 81.81% and XLM-R at 78.81%. However, mBERT exhibits lower precision, whereas LaBSE provides a better balance between precision (87.02%) and accuracy (86.46%), resulting in an F1-score of 84.33%. XLM-R demonstrates lower sensitivity but maintains competitive precision (85.47%) and accuracy (84.60%), with an F1-score of 82.00%. Sensitivity analysis based on the number of evidence reveals distinct model behaviors, where mBERT performs optimally with six pieces of evidence, XLM-R is more effective under limited evidence conditions, and LaBSE shows a stable and increasing sensitivity trend as the amount of evidence increases, indicating robustness in multi-evidence scenarios. Further statistical analysis shows that XLM-R has the lowest performance variance, while LaBSE statistically outperforms mBERT in several evaluation aspects. Overall, LaBSE is recommended as the most balanced model for multi-evidence-based political fact verification.


Keywords


Political News; Fact Verification; Multi-Evidence; Multilingual Model

Full Text:

PDF

References


A. Bucciol, “False claims in politics: Evidence from the US,” Research in Economics, vol. 72, no. 2, pp. 196–210, Jun. 2018, doi: 10.1016/j.rie.2018.04.002.

A. Barrón-Cedeño, I. Jaradat, G. Da San Martino, and P. Nakov, “Proppy: Organizing the news based on their propagandistic content,” Inf Process Manag, vol. 56, no. 5, pp. 1849–1864, Sep. 2019, doi: 10.1016/j.ipm.2019.03.005.

D. Calvo, L. Valera-Ordaz, M. Requena i Mora, and G. Llorca-Abad, “Fact-checking in Spain: Perception and trust,” Catalan Journal of Communication & Cultural Studies, vol. 14, no. 2, pp. 287–305, Oct. 2022, doi: 10.1386/cjcs_00073_1.

D. Fardiah, F. Darmawan, and R. Rinawati, “Fact-checking Literacy of Covid-19 Infodemic on Social Media in Indonesia,” Komunikator, vol. 14, no. 1, pp. 14–29, May 2022, doi: 10.18196/jkm.14459.

R. K. Kaliyar, A. Goswami, and P. Narang, “FakeBERT: Fake news detection in social media with a BERT-based deep learning approach,” Multimed Tools Appl, 2021, doi: 10.1007/s11042-020-10183-2.

M. Larki and E. Manouchehri, “Dispelling Fake News and Infodemic Management about COVID-19 Vaccination: A Literature Review NewsandInfodem,” Journal of Health Literacy, vol. 7, no. 3, pp. 91–105, 2022, doi: 10.22038/jhl.2022.65215.1289.

C. Chen, W. Chen, J. Zheng, A. Luo, F. Cai, and Y. Zhang, “Input-oriented demonstration learning for hybrid evidence fact verification,” Expert Syst Appl, vol. 246, p. 123191, Jul. 2024, doi: 10.1016/j.eswa.2024.123191.

H. Gong, C. Wang, and X. Huang, “Double Graph Attention Network Reasoning Method Based on Filtering and Program-Like Evidence for Table-Based Fact Verification,” IEEE Access, vol. 11, pp. 86859–86871, 2023, doi: 10.1109/ACCESS.2023.3304915.

Y. Zhu, J. Si, Y. Zhao, H. Zhu, D. Zhou, and Y. He, “EXPLAIN, EDIT, GENERATE: Rationale-Sensitive Counterfactual Data Augmentation for Multi-hop Fact Verification,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, Oct. 2023, pp. 13377–13392. doi: 10.18653/v1/2023.emnlp-main.826.

Y. Yang, Y. Zhou, Q. Ying, Z. Qian, and X. Zhang, “Search, Examine and Early-Termination: Fake News Detection with Annotation-Free Evidences,” Jul. 2024, [Online]. Available: http://arxiv.org/abs/2407.07931

W. Xu, J. Wu, Q. Liu, S. Wu, and L. Wang, “Evidence-aware Fake News Detection with Graph Neural Networks,” in WWW 2022 - Proceedings of the ACM Web Conference 2022, Association for Computing Machinery, Inc, Apr. 2022, pp. 2501–2510. doi: 10.1109/ICECA.2018.8474668.

N. Vo and K. Lee, “Hierarchical Multi-head Attentive Network for Evidence-aware Fake News Detection,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, Stroudsburg, PA, USA: Association for Computational Linguistics, Feb. 2021, pp. 965–975. doi: 10.18653/v1/2021.eacl-main.83.

L. Silva and L. Barbosa, “Improving dense retrieval models with LLM augmented data for dataset search,” Knowl Based Syst, vol. 294, p. 111740, Jun. 2024, doi: 10.1016/j.knosys.2024.111740.

X. Zhang and W. Gao, “Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method,” in Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Stroudsburg, PA, USA: Association for Computational Linguistics, Sep. 2023, pp. 996–1011. doi: 10.18653/v1/2023.ijcnlp-main.64.

N. Tiyajamorn, T. Kajiwara, Y. Arase, and M. Onizuka, “Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity Estimation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 7764–7774. doi: 10.18653/v1/2021.emnlp-main.612.

G. Z. Nabiilah, I. N. Alam, E. S. Purwanto, and M. F. Hidayat, “Indonesian multilabel classification using IndoBERT embedding and MBERT classification,” International Journal of Electrical and Computer Engineering, vol. 14, no. 1, pp. 1071–1078, Feb. 2024, doi: 10.11591/ijece.v14i1.pp1071-1078.

G. Li, Z. Wang, M. Zhao, Y. Song, and L. Lan, “Sentiment Analysis of Political Posts on Hong Kong Local Forums Using Fine-Tuned mBERT,” in Proceedings - 2022 IEEE International Conference on Big Data, Big Data 2022, Institute of Electrical and Electronics Engineers Inc., 2022, pp. 6763–6765. doi: 10.1109/BigData55660.2022.10020704.

A. Chauhan, T. Agrawal, and A. Singh, “Advancing Linguistic Frontiers with mBERT Fine-Tuning for Hindi and English Named Entity Recognition Using HiNER and WikiNEuRal Datasets,” in 2024 15th International Conference on Computing Communication and Networking Technologies, ICCCNT 2024, Institute of Electrical and Electronics Engineers Inc., 2024. doi: 10.1109/ICCCNT61001.2024.10724051.

N. Rathod, N. Mistry, D. Talati, M. Parikh, A. Kore, and P. Kanani, “Marathi Social Media Opinion Mining using XLM-R,” in Proceedings - International Conference on Applied Artificial Intelligence and Computing, ICAAIC 2022, Institute of Electrical and Electronics Engineers Inc., 2022, pp. 730–736. doi: 10.1109/ICAAIC53929.2022.9793308.

A. Gaurav, B. B. Gupta, S. Sharma, R. Bansal, and K. T. Chui, “XLM-RoBERTa Based Sentiment Analysis of Tweets on Metaverse and 6G,” Procedia Comput Sci, vol. 238, pp. 902–907, 2024, doi: 10.1016/j.procs.2024.06.110.

N. Rajapaksha, S. Ahangama, and S. Adikari, “Fine-tuning XLM-R for the Detection of Sinhala Hate Speech Content on Twitter and Youtube,” in ICARC 2023 - 3rd International Conference on Advanced Research in Computing: Digital Transformation for Sustainable Development, Institute of Electrical and Electronics Engineers Inc., 2023, pp. 19–23. doi: 10.1109/ICARC57651.2023.10145745.

G. Mehak, I. Muneer, and R. M. A. Nawab, “Urdu Text Reuse Detection at Phrasal level using Sentence Transformer-based approach,” Expert Syst Appl, vol. 234, p. 121063, Dec. 2023, doi: 10.1016/j.eswa.2023.121063.

H. Ma et al., “EX-FEVER: A Dataset for Multi-hop Explainable Fact Verification,” in Findings of the Association for Computational Linguistics ACL 2024, Stroudsburg, PA, USA: Association for Computational Linguistics, Oct. 2024, pp. 9340–9353. doi: 10.18653/v1/2024.findings-acl.556.

T. Pires, E. Schlinger, and D. Garrette, “How Multilingual is Multilingual BERT?,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4996–5001. doi: 10.18653/v1/P19-1493.

A. Conneau et al., “Unsupervised Cross-lingual Representation Learning at Scale,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 8440–8451. doi: 10.18653/v1/2020.acl-main.747.

F. Feng, Y. Yang, D. Cer, N. Arivazhagan, and W. Wang, “Language-agnostic BERT Sentence Embedding,” Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 878–891, Jul. 2022, doi: 10.18653/v1/2022.acl-long.62.

S. Shukla, H. Dutta, and P. Bhattacharyya, “Recon, Answer, Verify: Agents in Search of Truth,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, Stroudsburg, PA, USA: Association for Computational Linguistics, 2025, pp. 2429–2448. doi: 10.18653/v1/2025.emnlp-industry.167.

M. Soprano et al., “The many dimensions of truthfulness: Crowdsourcing misinformation assessments on a multidimensional scale,” Inf Process Manag, vol. 58, no. 6, p. 102710, Nov. 2021, doi: 10.1016/j.ipm.2021.102710.

M. Naseer, M. Asvial, and R. F. Sari, “An Empirical Comparison of BERT, RoBERTa, and Electra for Fact Verification,” in 3rd International Conference on Artificial Intelligence in Information and Communication, ICAIIC 2021, Institute of Electrical and Electronics Engineers Inc., Apr. 2021, pp. 241–246. doi: 10.1109/ICAIIC51459.2021.9415192.

Y. Li et al., “Deep learning-based platform performs high detection sensitivity of intracranial aneurysms in 3D brain TOF-MRA: An external clinical validation study,” Int J Med Inform, vol. 188, Aug. 2024, doi: 10.1016/j.ijmedinf.2024.105487.

S. Jamshidi et al., “Effective text classification using BERT, MTM LSTM, and DT,” Data Knowl Eng, vol. 151, May 2024, doi: 10.1016/j.datak.2024.102306.

E. A. U. Malahina, M. Saitakela, S. J. Bulan, M. I. J. Lamabelawa, and Y. S. Belutowe, “Teachable Machine: Optimization of Herbal Plant Image Classification Based on Epoch Value, Batch Size and Learning Rate,” Journal of Applied Data Sciences, vol. 5, no. 2, pp. 532–545, May 2024, doi: 10.47738/jads.v5i2.206.

S. Asy Syifa and I. Amelia Dewi, “MIND (Multimedia Artificial Intelligent Networking Database Arsitektur Resnet-152 dengan Perbandingan Optimizer Adam dan RMSProp untuk Mendeteksi Penyakit Paru-Paru,” Journal MIND Journal | ISSN, vol. 7, no. 2, pp. 139–150, 2022, doi: 10.26760/mindjournal.v7i2.139-150.

A. Gaurav, B. B. Gupta, S. Sharma, R. Bansal, and K. T. Chui, “XLM-RoBERTa Based Sentiment Analysis of Tweets on Metaverse and 6G,” Procedia Comput Sci, vol. 238, pp. 902–907, 2024, doi: 10.1016/j.procs.2024.06.110.




DOI: https://doi.org/10.47738/jads.v7i2.1198

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0