Detecting And Classifying Multi-Label Semantic Bias In 3,829 Indonesian Military Criminal Judgments Dataset Using Language Modeling And Ensemble Strategies
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.5372Keywords:
ensemble classification, Indonesia, language modeling, military criminal law, NLP, semantic biasAbstract
Objectivity in military criminal judgments is crucial for judicial legitimacy but is frequently compromised by semantic bias. To the best of our knowledge, this is the first study to specifically address automated bias detection within the Indonesian military legal domain, bridging a significant gap in the literature that has predominantly focused on general civil law. This study aims to develop a multi-label classification model to automatically detect and classify three specific types of bias (emotional, character, and ambiguity) in military legal texts. The methodology involved the acquisition and expert annotation of 3,829 judgment documents (2020–2025). Three feature extraction strategies (TF-IDF, IndoBERT, and Doc2Vec) were comparatively evaluated using KNN, MLP, Random Forest, and Custom Ensemble algorithms. Experimental results demonstrate that the lexical approach using Random Forest with TF-IDF achieved superior performance with a weighted F1-Score of 0.82, outperforming both complex embedding-based models and the ensemble approach (F1-Score 0.77). The findings further reveal that character bias is the most dominant form of distortion in the corpus. This research makes three novel contributions: (1) providing the first annotated legal dataset for the Indonesian military domain; (2) demonstrating the superior efficacy of lexical features (TF-IDF) over complex embeddings in this specific legal domain; and (3) establishing a technical foundation for a decision-support system to enhance judicial objectivity.
Downloads
References
C. I. Situmorang and I. Triadi, “Reformasi Kekuasaan Kehakiman di Indonesia: Meningkat, Independensi, dan Kualitas: (Judicial Power Reform in Indonesia: Improving Independence, Transparency, and Quality),” Journal Customary Law, vol. 1, no. 2, pp. 9–9, May 2024, doi: 10.47134/jcl.v1i2.2429.
C. T. Rahayu and I. Triadi, “Dualisme Peradilan Militer dan Peradilan Umum: Problematika dan Urgensi Reformasi (The Dualism Of Military And Civilian Courts: Issues And The Urgency Of Reform),” Media Hukum Indonesia (MHI), vol. 3, no. 3, Jun. 2025, doi: 10.5281/zenodo.15690834.
K. Javed and J. Li, “Artificial intelligence in judicial adjudication: Semantic biasness classification and identification in legal judgement (SBCILJ),” Heliyon, vol. 10, no. 9, p. e30184, May 2024, doi: 10.1016/j.heliyon.2024.e30184.
N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A Survey on Bias and Fairness in Machine Learning,” ACM Comput. Surv., vol. 54, no. 6, p. 115:1-115:35, Jul. 2021, doi: 10.1145/3457607.
A. Asista and R. A. Suntara, “Analisis Kesalahan Penggunaan Bahasa Indonesia dalam Laras Hukum pada Direktori Putusan Mahkamah Agung Republik Indonesia,” STILISTIKA, vol. 17, no. 1, pp. 69–82, Jan. 2024, doi: 10.30651/st.v17i1.20970.
N. F. P. Istiqomah, “Legitimasi Putusan Lembaga Adat dalam Sistem Hukum Nasional: Studi terhadap Peraturan Perundang-Undangan di Indonesia,” Jurnal Hukum Lex Generalis, vol. 6, no. 3, Mar. 2025, doi: 10.56370/jhlg.v6i3.840.
M. R. Maulana and P. Pujiyono, “Restrukturisasi Independensi Hakim dalam Sistem Peradilan Pidana yang Berwawasan Pancasila,” Jurnal Ilmiah Universitas Batanghari Jambi, vol. 21, no. 2, pp. 580–587, Jul. 2021, doi: 10.33087/jiubj.v21i2.1387.
H. B. Dogru, S. Tilki, A. Jamil, and A. Ali Hameed, “Deep Learning-Based Classification of News Texts Using Doc2Vec Model,” in 2021 1st International Conference on Artificial Intelligence and Data Analytics (CAIDA), Riyadh, Saudi Arabia: IEEE, Apr. 2021, pp. 91–96. doi: 10.1109/CAIDA51941.2021.9425290.
H. Zhong, C. Xiao, C. Tu, T. Zhang, Z. Liu, and M. Sun, “How Does NLP Benefit Legal System: A Summary of Legal Artificial Intelligence,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds., Online: Association for Computational Linguistics, Jul. 2020, pp. 5218–5230. doi: 10.18653/v1/2020.acl-main.466.
P. Joseph and S. Y. Yerima, “A comparative study of word embedding techniques for SMS spam detection,” in 2022 14th International Conference on Computational Intelligence and Communication Networks (CICN), Al-Khobar, Saudi Arabia: IEEE, Dec. 2022, pp. 149–155. doi: 10.1109/CICN56167.2022.10008245.
I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, “LEGAL-BERT: The Muppets straight out of Law School,” in Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu, Eds., Online: Association for Computational Linguistics, Nov. 2020, pp. 2898–2904. doi: 10.18653/v1/2020.findings-emnlp.261.
F. Ariai, J. Mackenzie, and G. Demartini, “Natural Language Processing for the Legal Domain: A Survey of Tasks, Datasets, Models, and Challenges,” ACM Comput. Surv., vol. 58, no. 6, p. 163:1-163:37, Dec. 2025, doi: 10.1145/3777009.
M. Dusi, N. Arici, A. Emilio Gerevini, L. Putelli, and I. Serina, “Discrimination Bias Detection Through Categorical Association in Pre-Trained Language Models,” IEEE Access, vol. 12, pp. 162651–162667, 2024, doi: 10.1109/ACCESS.2024.3482010.
C. Farrelly et al., “Current Topological and Machine Learning Applications for Bias Detection in Text,” preprint, arXiv:2311.13495, Nov. 2023, doi: 10.48550/arXiv.2311.13495.
S. Kalyan, S. K. Vangibhurathachhi, and K. Srinivasa, "Advanced natural language processing for legal document analysis," International Journal of Computer Engineering and Management (IJCEM), vol. 8, no. 2, May 2025.
W. A. Hidayat and V. R. S. Nastiti, “Perbandingan kinerja pre‑trained IndoBERT‑base dan IndoBERT‑lite pada klasifikasi sentimen ulasan TikTok Tokopedia Seller Center dengan model IndoBERT,” JSiI (Jurnal Sistem Informasi), vol. 11, no. 2, pp. 13–20, Sept. 2024, doi: 10.30656/jsii.v11i2.9168.
B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, K.-F. Wong, K. Knight, and H. Wu, Eds., Suzhou, China: Association for Computational Linguistics, Dec. 2020, pp. 843–857. doi: 10.18653/v1/2020.aacl-main.85.
M. D. Maulana and C. S. K. Aditya, “Perbandingan IndoBERT dan Bi-LSTM Dalam Mendeteksi Pelanggaran Undang-Undang ITE,” SINTECH (Science and Information Technology) Journal, vol. 8, no. 1, pp. 52–59, Apr. 2025, doi: 10.31598/sintechjournal.v8i1.1846.
C. Jocelynne, I. L. Wijayakusuma, and L. P. I. Harini, “Detection of Political Hoax News Using Fine-Tuning IndoBERT,” Journal of Applied Informatics and Computing, vol. 9, no. 2, pp. 354–360, Mar. 2025, doi: 10.30871/jaic.v9i2.8989.
I. O. Gallegos et al., “Bias and Fairness in Large Language Models: A Survey,” Computational Linguistics, vol. 50, no. 3, pp. 1097–1179, Sep. 2024, doi: 10.1162/coli_a_00524.
Z. Xie and T. Lukasiewicz, “An Empirical Analysis of Parameter-Efficient Methods for Debiasing Pre-Trained Language Models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds., Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 15730–15745. doi: 10.18653/v1/2023.acl-long.876.
M. Bozdag, N. Sevim, and A. Koç, “Measuring and Mitigating Gender Bias in Legal Contextualized Language Models,” ACM Trans. Knowl. Discov. Data, vol. 18, no. 4, pp. 1–26, May 2024, doi: 10.1145/3628602.
G. Proebsting and A. Poliak, “Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference,” in Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert, Eds., Abu Dhabi, UAE: Association for Computational Linguistics, Jan. 2025, pp. 5836–5851. Accessed: Dec. 20, 2025. [Online]. Available: https://aclanthology.org/2025.coling-main.389/
D. Tsarapatsanis and N. Aletras, “On the Ethical Limits of Natural Language Processing on Legal Text,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, C. Zong, F. Xia, W. Li, and R. Navigli, Eds., Online: Association for Computational Linguistics, Aug. 2021, pp. 3590–3599. doi: 10.18653/v1/2021.findings-acl.314.
H. Weerts, R. Xenidis, F. Tarissan, H. P. Olsen, and M. Pechenizkiy, “Algorithmic Unfairness through the Lens of EU Non-Discrimination Law: Or Why the Law is not a Decision Tree,” in Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, in FAccT ’23. New York, NY, USA: Association for Computing Machinery, June 2023, pp. 805–816. doi: 10.1145/3593013.3594044.
M. Kattnig, A. Angerschmid, T. Reichel, and R. Kern, “Assessing trustworthy AI: Technical and legal perspectives of fairness in AI,” Computer Law & Security Review, vol. 55, p. 106053, Nov. 2024, doi: 10.1016/j.clsr.2024.106053.
E. Akand, Y. Fan, W. Wobcke, S. A. Sisson, and M. S. Roque, “Diversity on the bench: An analysis of gendered biases in the language of Australian Family Law Court judgments,” PLOS ONE, vol. 20, no. 9, p. e0331841, Sep. 2025, doi: 10.1371/journal.pone.0331841.
J. V á z q u e z - O s o r i o et al., “LATE-GIL-NLP at SemEval-2025 Task 11: Multi-Language Emotion Detection and Intensity Classification Using Transformer Models with Optimized Loss Functions for Imbalanced Data,” in Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), S. Rosenthal, A. Rosá, D. Ghosh, and M. Zampieri, Eds., Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 666–674. Accessed: Dec. 20, 2025. [Online]. Available: https://aclanthology.org/2025.semeval-1.93/
K. Javed and J. Li, “Bias in adjudication: Investigating the impact of artificial intelligence, media, financial and legal institutions in pursuit of social justice,” PLOS ONE, vol. 20, no. 1, p. e0315270, Jan. 2025, doi: 10.1371/journal.pone.0315270.
H. Zakiri, A. H. Muhammad, and A. Nasiri, “Efektivitas Pelatihan Awal Berbasis Domain Spesifik Legal-BERT Untuk Natural Language Processing Hukum: Replikasi Dan Perluasan Studi Casehold,” Journal of Informatics, Electrical and Electronics Engineering, vol. 5, no. 1, pp. 9–20, Sep. 2025, doi: 10.47065/jieee.v5i1.2610.
S. R, K. T, P. K. P, and S. S. M, “Ensemble Text Classification with TF-IDF Vectorization for Hate Speech Detection in Social Media,” in 2023 International Conference on System, Computation, Automation and Networking (ICSCAN), PUDUCHERRY, India: IEEE, Nov. 2023, pp. 1–7. doi: 10.1109/ICSCAN58655.2023.10395354.
N. Nasution and S. Suprapto, “Court Decision Prediction Model Using Natural Language Processing and Random Forest,” IJCCS (Indonesian Journal of Computing and Cybernetics Systems), vol. 19, no. 3, pp. 305–316, Jul. 2025, doi: 10.22146/ijccs.108377.
C. R. Critcher, E. G. Helzer, and D. Tannenbaum, “Moral character evaluation: Testing another’s moral-cognitive machinery,” Journal of Experimental Social Psychology, vol. 87, p. 103906, Mar. 2020, doi: 10.1016/j.jesp.2019.103906.
A. Tantos and N. Amvrazis, “Classification and identification level ambiguity in error annotation,” Applied Corpus Linguistics, vol. 2, no. 3, p. 100035, Dec. 2022, doi: 10.1016/j.acorp.2022.100035.
S. Corbett-Davies, J. D. Gaebler, H. Nilforoshan, R. Shroff, and S. Goel, “The Measure and Mismeasure of Fairness,” Journal of Machine Learning Research, vol. 24, no. 312, pp. 1–117, 2023.
P. Radanliev, “AI Ethics: Integrating Transparency, Fairness, and Privacy in AI Development,” Applied Artificial Intelligence, vol. 39, no. 1, p. 2463722, Dec. 2025, doi: 10.1080/08839514.2025.2463722.
S. T.y.s.s and I. Chowdhury, “Fairness Beyond Performance: Revealing Reliability Disparities Across Groups in Legal NLP,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds., Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 24376–24390. doi: 10.18653/v1/2025.acl-long.1188.
L. Z. S. Sudar, J. L. Imbenay, I. Budi, A. Ramadiah, P. K. Putra, and A. B. Santoso, “Textual Analysis for Public Sentiment Toward National Police Using CRISP-DM Framework,” RIA, vol. 38, no. 1, pp. 63–72, Feb. 2024, doi: 10.18280/ria.380107.
E. Q. Nuranti, E. Yulianti, and H. S. Husin, “Predicting the Category and the Length of Punishment in Indonesian Courts Based on Previous Court Decision Documents,” Computers, vol. 11, no. 6, June 2022, doi: 10.3390/computers11060088.
N. Casali, B. E. Hilbig, and I. Thielmann, “A Critical Evaluation of the Factor Structure and Validity of the Moral Character Questionnaire,” Collabra: Psychology, vol. 11, no. 1, p. 138646, June 2025, doi: 10.1525/collabra.138646.
C. Ma and J. Lauwereyns, “Predictive cues elicit a liminal confirmation bias in the moral evaluation of real-world images,” Front. Psychol., vol. 15, Feb. 2024, doi: 10.3389/fpsyg.2024.1329116.
S. Y. Lee, R. Scheinberg, A. Shore, and A. Agrawal, “Who Relies More on World Knowledge and Bias for Syntactic Ambiguity Resolution: Humans or LLMs?,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang, Eds., Albuquerque, New Mexico: Association for Computational Linguistics, Apr. 2025, pp. 3484–3498. doi: 10.18653/v1/2025.naacl-long.177.
B. A. J. Jap and Y.-Y. Hsu, “An ERP study on verb bias and thematic role assignment in standard Indonesian,” Sci Rep, vol. 15, no. 1, p. 11847, Apr. 2025, doi: 10.1038/s41598-025-96240-y.
Y. Li, Y. Yang, P. Song, L. Duan, and R. Ren, “An improved SMOTE algorithm for enhanced imbalanced data classification by expanding sample generation space,” Sci Rep, vol. 15, no. 1, p. 23521, July 2025, doi: 10.1038/s41598-025-09506-w.
E. K. Tokpo, “Assessing and mitigating bias in natural language systems,” University of Antwerp, Antwerp, 2024. doi: 10.63028/10067/2090950151162165141.
R. I. Perwira, V. A. Permadi, D. I. Purnamasari, and R. P. Agusdin, “Domain-Specific Fine-Tuning of IndoBERT for Aspect-Based Sentiment Analysis in Indonesian Travel User-Generated Content,” Journal of Information Systems Engineering and Business Intelligence, vol. 11, no. 1, pp. 30–40, Mar. 2025, doi: 10.20473/jisebi.11.1.30-40.
A. S. Ekakristi, A. F. Wicaksono, and R. Mahendra, “Intermediate-task transfer learning for Indonesian NLP tasks,” Natural Language Processing Journal, vol. 12, p. 100161, Sep. 2025, doi: 10.1016/j.nlp.2025.100161.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Bayu Ardiyansyah, Lutfi Indra Nur Praditya, Galih Wasis Wicaksono, Nur Putri Hidayah

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



