Automatic Analysis of Political Discourse: A Comparative Study of Multilingual and Large Language Models

Ayaulym Sairanbekova, Aizhan Nazyrova, Gulmira Bekmanova, Lena Zhetkenbay, Banu Yergesh, Zhanar Lamasheva

Abstract


This paper proposes the growing importance of automated analysis of political discourse in low-resource languages, using the Kazakh language as a case study. As political communication in Kazakhstan has increasingly moved online between 2019 and 2023, the need for accurate tools to evaluate political sentiment has grown. However, limited linguistic resources in Kazakh have hindered tool development. This paper introduces the first annotated corpus of political discourse in Kazakh, comprising 3,022 sentences selected from official statements, televised debates, policy documents, and social media publications. Each text was manually annotated for political sentiment by expert linguists and political scientists, with inter-annotator agreement measured to confirm reliability. Two main methodological approaches were employed for automatic sentiment classification: adapting multilingual neural network models to the Kazakh corpus and testing advanced generative language models in scenarios with minimal training examples. Performance was evaluated using standard classification procedures. The inclusion of pragmatic features such as code-switching, rhetorical emphasis, and discursive context led to notable improvements in classification accuracy. Experimental results demonstrate that models adapted to multilingual input achieved high classification quality, with fine-tuned multilingual transformer models reaching F₁-scores of up to 0.90, while large language models reached an F₁-score of 0.94 in few-shot settings. Explicit modeling of code-switching and pragmatic features yielded an improvement of approximately 4 percentage points in F₁. This research contributes a practical resource and a methodological framework for analyzing political sentiment in underrepresented languages, highlighting the feasibility of developing high-quality automated tools for political text analysis without extensive training data.


Keywords


Political Sentiment Analysis; Pre-trained Language Models (PLMs); Large Language Models (LLMs); GPT-5; Gemini 2.5 Pro; Zero-shot Learning; Few-shot Learning; Code-switching; Low-resource Language; Kazakh Language

Full Text:

PDF

References


M. Ansari, M. B. Aziz, M. O. Siddiqui, H. Mehra, and K. P. Singh, “Analysis of political sentiment orientations on Twitter,” Procedia Computer Science, vol. 167, pp. 1821–1828, 2020. [Online]. Available: https://doi.org/10.1016/j.procs.2020.03.201

A. Bhattacharjee et al., “BanglaBERT: Language model pretraining and benchmarks for low-resource language understanding evaluation in Bangla,” arXiv preprint arXiv:2101.00204, 2021. [Online]. Available: https://doi.org/10.48550/arXiv.2101.00204

J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186. [Online]. Available: https://doi.org/10.18653/v1/N19-1423

T. Kaufmann, P. Weng, V. Bengs, and E. Hüllermeier, “A survey of reinforcement learning from human feedback,” 2024. [Online]. Available: https://epub.ub.uni-muenchen.de/125328/1/2312.14925v2.pdf

X. Deng, V. Bashlovkina, F. Han, S. Baumgartner, and M. Bendersky, “What do LLMs know about financial markets? A case study on Reddit market sentiment analysis,” in Companion Proc. ACM Web Conf. 2023, Austin, TX, USA, pp. 107–110. [Online]. Available: https://doi.org/10.1145/3543873.3587324

Z. Nasim, Q. Rajput, and S. Haider, “Sentiment analysis of student feedback using machine learning and lexicon-based approaches,” in Proc. 2017 Int. Conf. Research and Innovation in Information Systems (ICRIIS), Langkawi, Malaysia, 2017, pp. 1–6. [Online]. Available: https://doi.org/10.1109/ICRIIS.2017.8002475

M. Loukili, F. Messaoudi, and M. El Ghazi, “Sentiment analysis of product reviews for e-commerce recommendation based on machine learning,” Int. J. Advances in Soft Computing & Its Applications, vol. 15, no. 1, pp. 1–13, 2023. [Online]. Available: https://doi.org/10.15849/IJASCA.230320.01

E. R. Rhythm, R. Shuvo, M. S. Hossain, M. Islam, and A. A. Rasel, “Sentiment analysis of restaurant reviews from Bangladeshi food delivery apps,” in Proc. 2023 Int. Conf. Emerging Smart Computing and Informatics (ESCI), Pune, India, 2023. [Online]. Available: https://doi.org/10.1109/ESCI56872.2023.10100214

K. Rahul, B. Jindal, K. Singh, and P. Meel, “Analysing public sentiments regarding COVID-19 vaccine on Twitter,” in Proc. 2021 7th Int. Conf. Advanced Computing and Communication Systems (ICACCS), Coimbatore, India, 2021, pp. 488–493. [Online]. Available: https://doi.org/10.1109/ICACCS51430.2021.9441693

M. Hoq, P. Haque, and M. N. Uddin, “Sentiment analysis of Bangla language using deep learning approaches,” in Int. Conf. Computing Science, Communication and Security, Cham, Switzerland: Springer, 2021, pp. 140–151. [Online]. Available: https://doi.org/10.1007/978-3-030-76776-1_10

S. Sarker, “BanglaBERT: Bengali mask language model for Bengali language understanding,” GitHub Repository, 2020. [Online]. Available: https://github.com/sagorbrur/bangla-bert. [Accessed: Oct. 20, 2025].

A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised cross-lingual representation learning at scale,” arXiv preprint arXiv:1911.02116, 2019, doi: 10.48550/arXiv.1911.02116.

F. T. J. Faria, M. B. Moin, R. I. Mumu, M. M. A. Abir, A. N. Alfy, and M. S. Alam, “Motamot: A dataset for revealing the supremacy of large language models over transformer models in Bengali political sentiment analysis,” in Proc. IEEE Region 10 Symposium (TENSYMP), 2024, pp. 1–8.

B. Pahwa and B. Pahwa, “BpHigh at SemEval-2023 Task 7: Can fine-tuned cross-encoders outperform GPT-3.5 in NLI tasks on clinical trial data?,” in Proc. 17th Int. Workshop on Semantic Evaluation (SemEval-2023), Toronto, Canada, 2023, pp. 1936–1944. [Online]. Available: https://aclanthology.org/2023.semeval-1.266/

J. Ye, X. Chen, N. Xu, C. Zu, Z. Shao, S. Liu, and X. Huang, “Comprehensive capability analysis of GPT-3 and GPT-3.5 series models,” arXiv preprint arXiv:2303.10420, 2023, doi: 10.48550/arXiv.2303.10420.

F. T. J. Faria, M. B. Moin, R. I. Mumu, M. M. A. Abir, A. N. Alfy, and M. S. Alam, “Motamot: A dataset for revealing the supremacy of large language models over transformer models in Bengali political sentiment analysis,” arXiv e-prints, arXiv:2407, 2024.

R. Faliotco and P. Quatto, “Fleiss’ Kappa statistic without paradoxes,” Quality & Quantity, vol. 49, no. 2, pp. 463–470, 2014, doi: 10.1007/s11135-014-0003-1.

R. Yeshpanov and H. A. Varol, “KazSAnDRA: Kazakh sentiment analysis dataset of reviews and attitudes,” in Proc. LREC 2024, Torino, Italy, 2024. [Online]. Available: https://aclanthology.org/2024.lrec-main.844

S. Akhmedov and A. Nugumanova, “Development of sentiment analysis model in Kazakh language to analyze reviews,” Preprints, 2024. [Online]. Available: https://www.preprints.org/manuscript/202405.1300/v1

Hugging Face, “RemBERT Sentiment Analysis Polarity Classification (Kazakh),” Model Card, 2024. [Online]. Available: https://huggingface.co/issai/rembert-sentiment-analysis-polarity-classification-kazakh [Accessed: Oct. 16, 2025].

M. Goloburda et al., “Qorgau: Evaluating LLM safety in Kazakh-Russian bilingual contexts,” arXiv preprint arXiv:2502.13640, 2025, doi: 10.48550/arXiv.2502.13640.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186.

H. W. Chung, T. Fevry, H. Tsai, M. Johnson, and S. Ruder, “Rethinking embedding coupling in pre-trained language models,” arXiv preprint arXiv:2010.12821, 2020, doi: 10.48550/arXiv.2010.12821.




DOI: https://doi.org/10.47738/jads.v7i2.1118

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0