Comparing Pre-Norm and Post-Norm Transformers in Preserving Gender Information for Indonesian–English Translation through Attention-Based Signal Reinforcement

Andik Wijanarko, Rinaldi Munir, Masayu Leylia Khodra, Dessi Puji Lestari

Abstract


Gender realization in Indonesian–English machine translation remains challenging due to the absence of grammatical gender in Indonesian, which often leads to unstable or ambiguous gender representations in English outputs. While Transformer-based models have demonstrated strong general translation performance, their ability to preserve gender information across encoding layers remains inconsistent and poorly understood, particularly with respect to architectural normalization strategies.

This study presents a comparative analysis of Pre-Norm and Post-Norm Transformer architectures in preserving gender information, and examines the role of attention-based signal reinforcement in mitigating representational degradation. The reinforcement mechanism is introduced prior to standard encoder processing to strengthen gender-relevant token interactions without modifying the overall model structure.

Four controlled configurations—Post-Norm, Pre-Norm, Post-Norm with attention-based reinforcement, and Pre-Norm with attention-based reinforcement—are trained under identical random seeds on both unbalanced and balanced datasets. Evaluation is performed on gender-ambiguous test sentences without explicit gender annotations to assess generalization. Gender preservation is assessed at the output level using gender-specific accuracy and BLEU score, and at the representation level using cosine similarity between gender cue embeddings and English gendered pronouns.

The results show that Post-Norm Transformers fail to maintain stable gender representations, yielding near-random gender accuracy (~50%) and negligible BLEU scores. Pre-Norm architectures improve training stability but achieve limited gender accuracy (around 30%). Incorporating attention-based signal reinforcement substantially enhances gender preservation, with accuracy rising to over 50% and reaching up to 56% under balanced training conditions, accompanied by a consistent increase in cosine similarity values (exceeding 0.35) between gender cues and corresponding pronouns. These findings indicate that normalization strategy and attention-based reinforcement jointly determine the stability of gender representations in Transformer-based machine translation.


Keywords


Pre-Norm and Post-Norm Transformers, Signal Reinforcement, Indonesian–English Translation, Gender Representation Stability

Full Text:

PDF

References


S. K. Mondal, H. Zhang, H. M. D. Kabir, K. Ni, dan H.-N. Dai, “Machine translation and its evaluation: a study,” Artif. Intell. Rev., vol. 56, no. 9, hal. 10137–10226, 2023, doi: 10.1007/s10462-023-10423-5.

M. Salıcı dan Ü. E. Ölçer, “Impact of Transformer-Based Models in NLP: An In-Depth Study on BERT and GPT,” in 2024 8th International Artificial Intelligence and Data Processing Symposium (IDAP), 2024, hal. 1–6. doi: 10.1109/IDAP64064.2024.10710796.

L. Kang, S. He, M. Wang, F. Long, dan J. Su, “Bilingual attention based neural machine translation,” Appl. Intell., vol. 53, no. 4, hal. 4302–4315, 2023, doi: 10.1007/s10489-022-03563-8.

Y. B. Kaya dan A. C. Tantuğ, “Effect of tokenization granularity for Turkish large language models,” Intell. Syst. with Appl., vol. 21, hal. 200335, 2024, doi: https://doi.org/10.1016/j.iswa.2024.200335.

M. C. Roy, S. K. Bisoy, dan P. K. Das, “A Diffusion Driven Multimodal Fusion Framework for Context Aware Sarcasm Detection via Sentiment Syntax Graph Modeling,” Arab. J. Sci. Eng., 2025, doi: 10.1007/s13369-025-10848-w.

U. Aut, “Large Language Models ‘ Ad Referendum ’: How Good are They at Machine Translation in The Legal Domain,” MONTI, vol. 16, hal. 75–107, 2024.

C.-O. Truică, A.-I. Stan, dan E.-S. Apostol, “SimpLex: a lexical text simplification architecture,” Neural Comput. Appl., vol. 35, no. 8, hal. 6265–6280, 2023, doi: 10.1007/s00521-022-07905-y.

F. Algobaei, E. Alzain, E. Naji, dan K. A. Nagi, “Gender Issues between Gemini and ChatGPT: The Case of English-Arabic Translation,” World J. English Lang., vol. 15, no. 1, hal. 9, 2024, doi: 10.5430/wjel.v15n1p9.

J. D. Id dan W. K. Id, “Beyond the spotlight : Unveiling the gender bias curtain in movie reviews,” hal. 1–22, 2025, doi: 10.1371/journal.pone.0316093.

Z. M. Nia, A. Ahmadi, B. Mellado, J. Wu, dan J. Orbinski, “Twitter-based gender recognition using transformers,” Math. Biosci. Eng., vol. 20, no. October 2022, hal. 15962–15981, 2023, doi: 10.3934/mbe.2023711.

A. Vellintihun, T. Doren, dan K. Amitab, “Neural machine translation systems for English to Khasi : A case study of an Austroasiatic language,” Expert Syst. Appl., vol. 238, no. PA, hal. 121813, 2024, doi: 10.1016/j.eswa.2023.121813.

S. Das dan J. H. Paik, “Context-sensitive gender inference of named entities in text,” Inf. Process. Manag., vol. 58, no. 1, hal. 102423, 2021, doi: https://doi.org/10.1016/j.ipm.2020.102423.

Nurtamin, H. Abbas, E. Iswary, dan M. Hasyim, “Gender Bias In Machine Translation ( Google Translate ) From Indonesian To English,” J. Posit. Sch. Psychol., vol. 6, no. 4, hal. 9754–9761, 2022.

B. Paul, “Multimodal Machine Translation Approaches for Indian Languages : A Comprehensive Survey,” J. ofUniversal Comput. Sci., vol. 30, no. 5, hal. 694–717, 2024, doi: 10.3897/jucs.109227.

S. Ahmad dan A. Maaytah, “Evaluating Three Neural Machine Translation Platforms for English-Arabic Translation : A Comparative Study of Linguistic Accuracy and Cultural Fidelity,” World J. English Lang., vol. 16, no. 2, hal. 1–14, 2026, doi: 10.5430/wjel.v16n2p1.

C. Valentina Bastías dan J. F. Medone, “Bilingual translation analysis of the existence of gender bias in the machine translations of DeepL and Google Translate,” Synerg. Chili, no. 19, hal. 37–56, 2023, [Daring]. Tersedia pada: https://www.scopus.com/inward/record.uri?eid=2-s2.0-85215208741&partnerID=40&md5=a5cce16b730ef5f3af8a8d0f38d4db54

Z. Y. Li Nguyen, Oliver Mayeux, “Code-switching input for machine translation, a case study of Vietnamese–English data.pdf,” Int. J. Multilinualism, vol. 21, no. 4, hal. 1–22, 2024.

Q. Jie dan C. Huang, “A Dual-Path Gated Attention-Based Deep Learning Model for Automated Essay Scoring Using Linguistic Features,” Int. J. Adv. Comput. Sci. Appl., vol. 16, no. 7, hal. 776–788, 2025.

M. Lardelli dan M. Lardelli, “Gender-fair translation : a case study beyond the binary,” Perspectives (Montclair)., vol. 6623, no. 32, 2024, doi: 10.1080/0907676X.2023.2268654.

N. Costa, D. Anseán, M. Dubarry, dan L. Sánchez, “ICFormer: A Deep Learning model for informed lithium-ion battery diagnosis and early knee detection,” J. Power Sources, vol. 592, hal. 233910, 2024, doi: https://doi.org/10.1016/j.jpowsour.2023.233910.

C. Bosco, V. Patti, S. Frenda, A. T. Cignarella, M. Paciello, dan F. D’Errico, “Detecting racial stereotypes: An Italian social media corpus where psychology meets NLP,” Inf. Process. Manag., vol. 60, no. 1, hal. 103118, 2023, doi: https://doi.org/10.1016/j.ipm.2022.103118.

S. N. Muralikrishna, R. Holla, N. Harivinod, dan R. Ganiga, “Cross-Lingual Short-Text Semantic Similarity for Kannada–English Language Pair,” Computers, vol. 13, no. 9, hal. 1–15, 2024, doi: 10.3390/computers13090236.

S. Crossley dan L. Holmes, “Assessing receptive vocabulary using state‑of‑the‑art natural language processing techniques,” J. Second Lang. Stud., vol. 6, no. 1, hal. 1–28, 2023, doi: 10.1075/jsls.22006.cro.

M. A. Ibrahim, “Prompt-Based Data Augmentation with Large Language Models for Indonesian Gender-Based Hate Speech Detection,” J. Comput. Sci. Orig., vol. 28, no. 8, 2024, doi: 10.3844/jcssp.2024.819.826.

S. Seli, “Cultural Influence on The Translation of ‘ Eclipse ’ Novel By Stephenie Meyer : A Semiotic Perspective,” Fonseca, J. Commun., hal. 280–293, 2024, doi: 10.48047/fjc.28.01.

Z. Arifin, A. Sunanda, A. H. Prabawa, dan A. Sabardila, “Equivalency, readability, and acceptability of information technology terms’ translation from english to Indonesian,” Int. J. Innov. Creat. Chang., vol. 12, no. 2, hal. 185–202, 2020, [Daring]. Tersedia pada: https://www.scopus.com/inward/record.uri?eid=2-s2.0-85083062539&partnerID=40&md5=8697ce9b3b268b5bc1f7aa67b22232b0

D. Guo, “Deep learning-driven context-aware English translation for ambiguous sentences Deep learning-driven context-aware English translation for ambiguous sentences Donghui Guo,” Int. J. Inf. Commun. Technol., vol. 26, no. March, 2025, doi: 10.1504/IJICT.2025.10071311.

Y. Tian, “Decoding social group representation in American literature using contextualized embedding analysis and bias detection algorithms,” J. Comput. Methods Sci. Eng., vol. 0, no. 0, hal. 14727978251393472, 2025, doi: 10.1177/14727978251393473.

G. Ji, Z. Chen, H. Liu, T. Liu, dan B. Wang, “applied sciences APTrans : Transformer-Based Multilayer Semantic and Locational Feature Integration for Efficient Text Classification,” Applid Sci., vol. 14, 2024.

L. V. Rodríguez, J. De La Rosa Yacomelo, R. R. González, dan D. G. Segura, “Pronouns in Wayuunaiki and Spanish: A Contrastive Analysis; [Pronoms en Wayuunaiki et Espagnol: Une analyse contrastive]; [Pronomes em wayuunaiki e espanhol: uma análise contrastiva]; [Pronombres en wayuunaiki y español: una mirada contrastiva],” Ikala, vol. 27, no. 1, hal. 153 – 173, 2022, doi: 10.17533/udea.ikala.v27n1a08.

C. Manna, A. Alishahi, dan E. Vanmassenhove, “Are We Paying Attention to Her ? Investigating Gender Disambiguation and Attention in Machine Translation,” in Proceedings ofthe 3rd Workshop on Gender-Inclusive Translation Technologies (GITT 2025), 2025, hal. 1–16.

D. Saunders, R. Sallis, dan B. Byrne, “Neural Machine Translation Doesn’t Translate Gender Coreference Right Unless You Make It,” Proc. ofthe Second Work. Gend. Bias Nat. Lang. Process. Barcelona, Spain (Online), December 13, 2020., hal. 35–43, 2020, [Daring]. Tersedia pada: http://arxiv.org/abs/2010.05332

M. Bernagozzi, B. Srivastava, F. Rossi, dan S. Usmani, “Gender Bias in Online Language Translators: Visualization, Human Perception, and Bias/Accuracy Tradeoffs,” IEEE Internet Comput., vol. 25, no. 5, hal. 53–63, 2021, doi: 10.1109/MIC.2021.3097604.

A. Qorbani, R. Ramezani, A. Baraani, dan A. Kazemi, “Neurocomputing Multilingual neural machine translation for low-resource languages by twinning important nodes,” Neurocomputing, vol. 634, no. March, hal. 129890, 2025, doi: 10.1016/j.neucom.2025.129890.

N. B. Rajaboina dan T. P. Sariki, “Enhanced subtitle generation in videos: leveraging hybrid BERT–CNN–LSTM architecture for contextual understanding,” Int. J. Inf. Technol., vol. 17, no. 7, hal. 3947–3953, 2025, doi: 10.1007/s41870-025-02610-0.

B. Gra, “Pragmatic Uses of Gestures in Brazilian Portuguese in Contexts of Negation,” Rev. Estud. da Ling., vol. 32, no. 2, hal. 560–577, 2024, doi: 10.17851/2237-2083.32.2.560.

S. Hu et al., “Neural Machine Translation by Fusing Key Information of Text,” Comput. Mater. Contin., 2023, doi: 10.32604/cmc.2023.032732.

X. Su, X. Zhao, J. Ren, Y. Li, dan M. Rätsch, “Pre-training neural machine translation with alignment information via optimal transport,” Multimed. Tools Appl., vol. 83, no. 16, hal. 48377–48397, 2024, doi: 10.1007/s11042-023-17479-z.

H. H. Vu, “Context-Aware Machine Translation with Source Coreference Explanation,” Trans. Assoc. Comput. Linguist., vol. 12, hal. 856–874, 2024.




DOI: https://doi.org/10.47738/jads.v7i2.1257

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0