Optimizing Function-Level Source Code Classification Using Meta-Trained CodeBERT in Low-Resource Settings

Abednego Dwi Septiadi, Muhamad Awiet Wiedanto Prasetyo, Geusan Edurais Aria Daffa

Abstract


This study investigates the effectiveness of a meta-trained transformer-based model, CodeBERT, for classifying source code functions in environments with limited labeled data. The primary objective is to improve the accuracy and generalizability of function-level code classification using few-shot learning, a strategy where the model learns from only a few labeled examples per category. We introduce a meta-learning framework designed to enable CodeBERT to adapt to new function types with minimal supervision, addressing a common limitation in traditional code classification methods that require extensive labeled datasets and manual feature engineering. The methodology involves episodic few-shot classification, where each episode simulates a low-resource task using five labeled and five unlabeled samples per function class. A balanced subset of Python functions was sampled from the CodeXGLUE benchmark, consisting of ten function categories with equal representation. The source code was preprocessed by removing comments and docstrings, then tokenized into a fixed length of 128 tokens to fit the model input format. The meta-trained CodeBERT was evaluated across 10 episodes, each representing a different task composition. Results show that the model achieves an average classification accuracy of 73.0%, with high accuracy on function categories characterized by unique syntax patterns, and lower performance on categories with overlapping logic or naming structures. Despite this variability, the model-maintained accuracy above 60% in all episodes. These findings suggest that meta-learning significantly enhances the adaptability of CodeBERT to unseen tasks under data-constrained conditions. This research demonstrates that meta-trained transformer models can serve as practical tools for real-time code analysis, particularly in integrated development environments and continuous integration pipelines. Future work may include extending the framework to other programming languages and incorporating semantic code representations to further reduce classification ambiguity.

Keywords


Meta-Learning; Few-Shot Learning; Code Classification; CodeBERT; Transformer Models; Low-Resource Software Engineering; Software Process

Full Text:

PDF

References


I. Bacher, "Visualising the Complex Features of Source Code," 2019. doi: 10.21427/D1AV-KS51.

F. M. Medeiros, G. Lima, G. Amaral, S. Apel, C. Kästner, M. Ribeiro, and R. Gheyi, "An Investigation of Misunderstanding Code Patterns in C Open-Source Software Projects," Empirical Software Engineering, vol. 24, no. 3, pp. 1693-1726, 2018. [Online]. doi: 10.1007/s10664-018-9666-x.

D. Bamidis, I. Kalouptsoglou, A. Ampatzoglou, and A. Chatzigeorgiou, "Software Skills Identification: A Multi-Class Classification on Source Code Using Machine Learning," Global Clinical Engineering Journal, 2024. [Online]. doi: 10.31354/globalce.v6isi6.278.

F. Fontana and M. Zanoni, "Code smell severity classification using machine learning techniques," Knowledge-Based Systems, vol. 128, pp. 43-58, 2017. [Online]. doi: 10.1016/J.KNOSYS.2017.04.014.

P. Singh et al., "Shifting to machine supervision: annotation-efficient semi and self-supervised learning for automatic medical image segmentation and classification," Journal of Medical Image Computing and Computer-Assisted Intervention (MICCAI), vol. 27, no. 3, pp. 1-12, 2023.

N. G. Laleh, H. Muti, C. Loeffler, et al., "Benchmarking weakly-supervised deep learning pipelines for whole slide classification in computational pathology," Medical Image Analysis, vol. 79, p. 102474, 2022.

L. Yan, Y. Zheng, and J. Cao, "Few-shot learning for short text classification," Multimedia Tools and Applications, vol. 77, no. 24, pp. 29799–29810, 2018.

M. Pan, H. Xin, C. Xia, and H. Shen, "Few-shot classification with task-adaptive semantic feature learning," Pattern Recognition, vol. 141, p. 109594, 2023.

S. Manghat, S. Manghat, and T. Schultz, "Few-shot meta multilabel classifier for low resource accented code-switched speech," 2023 26th Conference of the Oriental COCOSDA International Committee, 2023, pp. 1-6.

Y. Chen, Z. Liu, H. Xu, T. Darrell, and X. Wang, "Meta-Baseline: Exploring Simple Meta-Learning for Few-Shot Learning," 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 9042-9051.

X. Ma, C. Yu, X. Yang, and X. Chen, "Few-shot learning based on attention relation compare network," 2019 International Conference on Data Mining Workshops (ICDMW), 2019, pp. 658-664.

F. Teng, Q. Zhang, X. Zhou, J. Hu, and T. Li, "Few-shot ICD coding with knowledge transfer and evidence representation," Expert Systems with Applications, vol. 238, p. 121861, 2024.

A. S. Tripathi, M. Danelljan, L. Van Gool, and R. Timofte, "Fast Few-Shot Classification by Few-Iteration Meta-Learning," 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021, pp. 9522-9528.

C. Watson, "Deep Learning in Software Engineering," Dissertation, College of William & Mary, 2020. doi: 10.21220/S2-3BN1-FV48.

M. White, "Deep Representations for Software Engineering," 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering, vol. 2, pp. 781-783, 2015. doi: 10.1109/ICSE.2015.248.

M. Fan, A. Jia, J. Liu, T. Liu, and W. Chen, "When Representation Learning Meets Software Analysis," Proceedings of the 1st ACM SIGSOFT International Workshop on Representation Learning for Software Engineering and Program Languages, 2020.. doi: 10.1145/3416506.3423578.

D. Wang, W. Dong, and S. Li, "A Multi-Task Representation Learning Approach for Source Code," Proceedings of the 1st ACM SIGSOFT International Workshop on Representation Learning for Software Engineering and Program Languages, 2020. doi: 10.1145/3416506.3423575.

H. Samoaa, F. Bayram, P. Salza, and P. Leitner, "A Systematic Mapping Study of Source Code Representation for Deep Learning in Software Engineering," IET Software, vol. 16, pp. 351-385, 2022. doi: 10.1049/sfw2.12064.

F. Zhang, B. Chen, Y. Zhao, and X. Peng, "Slice-Based Code Change Representation Learning," 2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), 2023, pp. 319-330. doi: 10.1109/SANER56733.2023.00038.

C. Zan, L. Ding, L. Shen, Y. Cao, and W. Liu, "Code-Switching Finetuning: Bridging Multilingual Pretrained Language Models for Enhanced Cross-Lingual Performance," Engineering Applications of Artificial Intelligence, vol. 139, p. 109532, 2025. doi: 10.1016/j.engappai.2024.109532.

Y. Meng, J. Huang, Y. Zhang, Y. Zhang, and J. Han, "Pretrained Language Representations for Text Understanding: A Weakly-Supervised Perspective," Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2023. doi: 10.1145/3580305.3599569.

T. Yu, X. Gu, and B. Shen, "Code Question Answering via Task-Adaptive Sequence-to-Sequence Pre-training," 2022 29th Asia-Pacific Software Engineering Conference (APSEC), 2022, pp. 229-238. doi: 10.1109/APSEC57359.2022.00035.

K. Ma, S. Zheng, M. Tian, et al., "CnGeoPLM: Contextual Knowledge Selection and Embedding with Pretrained Language Representation Model for the Geoscience Domain," Earth Science Informatics, vol. 16, pp. 3629-3646, 2023. doi: 10.1007/s12145-023-01112-6.

G. Colavito, F. Lanubile, and N. Novielli, "Few-Shot Learning for Issue Report Classification," 2023 IEEE/ACM 2nd International Workshop on Natural Language-Based Software Engineering (NLBSE), 2023, pp. 16-19. doi: 10.1109/NLBSE59153.2023.00011.

R. Zhang, "Recent Advancement for Few-Shot Learning," Journal of Progress in Engineering and Physical Science, 2023. doi: 10.56397/jpeps.2023.12.06.

R. K. Helmeczi, M. Cevik, and S. Yıldırım, "Few-shot learning for sentence pair classification and its applications in software engineering," ArXiv, 2023.

V. Satorras and J. Bruna, "Few-Shot Learning with Graph Neural Networks," 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021.

L. Logeswaran, A. Lee, M. Ott, H. Lee, M. Ranzato, and A. D. Szlam, "Few-shot Sequence Learning with Transformers," 2020 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2020.

Y. Xie, H. Wang, B. Yu, and C. Zhang, "Secure collaborative few-shot learning," Knowledge-Based Systems, vol. 203, p. 106157, 2020. doi: 10.1016/j.knosys.2020.106157.

D. Zhan and H. Ye, "Few-shot learning via model composition," Science China Information Sciences, vol. 50, pp. 662-674, 2020. doi: 10.1360/n112018-00332.

R. Li, O. Bohdal, R. K. Mishra, H. Kim, D. Li, N. Lane, and T. M. Hospedales, "A Channel Coding Benchmark for Meta-Learning," 2021 International Workshop on Machine Learning (ML), 2021.

S. Manghat, S. Manghat, and T. Schultz, "Few-shot meta multilabel classifier for low resource accented code-switched speech," 2023 26th Conference of the Oriental COCOSDA International Committee, 2023, pp. 1-6. doi: 10.1109/o-cocosda60357.2023.10482982.

Y. Lee, W. Kim, and S. Choi, "Discrete Infomax Codes for Meta-Learning," Proceedings of the 2020 International Conference on Machine Learning (ICML), 2020.

N. Farajzadeh, G. Pan, Z. Wu, and M. Yao, "Multiclass Classification Based on Meta Probability Codes," International Journal of Pattern Recognition and Artificial Intelligence, vol. 25, pp. 1219-1241, 2011. doi: 10.1142/S021800141100910X.

F. Ciompi, O. Pujol, and P. Radeva, "A Meta-Learning Approach to Conditional Random Fields Using Error-Correcting Output Codes," 2010 20th International Conference on Pattern Recognition (ICPR), 2010, pp. 710-713. doi: 10.1109/ICPR.2010.179.

S. Wang, L. Sun, and J. Fang, "Molecular cancer classification using a meta-sample-based regularized robust coding method," BMC Bioinformatics, vol. 15, p. S2, 2014. doi: 10.1186/1471-2105-15-S15-S2.

X. Pan, F. Li, and L. Liu, "Improving Meta-Learning Classification with Prototype Enhancement," 2021 International Joint Conference on Neural Networks (IJCNN), 2021, pp. 1-8. doi: 10.1109/IJCNN52387.2021.9534292.

L. Fang, Z. Huang, Y. Zhou, and T. Chen, "Adaptive Code Completion with Meta-learning," Proceedings of the 12th Asia-Pacific Symposium on Internetware, 2020. doi: 10.1145/3457913.3457933.

X. Zheng, M. Jiang, and Z. Q. Zhou, "Boosting Metamorphic Relation Prediction via Code Representation Learning: An Empirical Study," Software Testing, Verification and Reliability, vol. 34, 2024. doi: 10.1002/stvr.1889.




DOI: https://doi.org/10.47738/jads.v6i3.902

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0