Implementation of Position Weighted Retrieval BM25 (PoWeR-BM25) for Thesis Recommendation Systems
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.5414Keywords:
BM25, Digital Library, Positional Weighting, Recommendation System, Thesis RetrievalAbstract
The increasing volume of academic documents in University of Udayana Digital Library (e-Perpus UNUD) has increased the difficulty for students to efficiently identify relevant prior research aligned with their research topics. Conventional keyword-based search mechanisms often fail to capture contextual relevance, resulting in information overload and suboptimal retrieval quality. To address this challenge, this study presents the design and implementation of a thesis recommendation system based on a modified BM25 algorithm called Position-Weighted Retrieval BM25 (PoWeR-BM25). The proposed model introduces an adaptive positional weighting scheme that emphasizes query term occurrences appearing at the beginning and end of a document, where essential contextual information is typically concentrated. The system was developed using an information retrieval framework and evaluated across multiple configurations combining different data representations (title, abstract, and title–abstract) and preprocessing techniques (stemming and non-stemming). Experimental results demonstrate that POWER-BM25 consistently outperforms traditional BM25 and TF-IDF in terms of Mean Average Precision (MAP) and F1-Score, particularly when applied to stemmed abstract data. The best performing configuration was subsequently deployed as recommendation feature integrated into University of Udayana Digital Library via a Web Service API, enabling users to obtain more accurate and contextually relevant thesis recommendations. These findings highlight the practical importance of incorporating positional term weighting into probabilistic retrieval models as a lightweight yet effective approach to improving academic recommendation systems without relying on computationally intensive machine learning models.
Downloads
References
S. Delliana, “Menggali Peran Perpustakaan Perguruan Tinggi di Zaman Digital: Systematic Literature Review,” Media Pustak., vol. 31, no. 3, pp. 318–330, Jan. 2025, doi: 10.37014/medpus.v31i3.5170.
C. Manning, P. Raghavan, and H. Schuetze, Introduction to Information Retrieval. Cambridge, England: Cambridge University Press, 2009.
I. B. N. W. Manuaba, G. R. Dantes, and G. Indrawan, “Analisis Sentimen Data Provider Layanan Internet Pada Twitter Menggunakan Support Vector Machine Dengan Penambahan Algoritma Levenshtein Distance,” J. SISKOM-KB Sist. Komput. Dan Kecerdasan Buatan, vol. 5, no. 2, pp. 9–17, Mar. 2022, doi: 10.47970/siskom-kb.v5i2.261.
P. Sidik, I. M. G. Sunarya, and I. G. A. Gunadi, “Comparison of Random Forest and Support Vector Machine Methods in Sentiment Analysis of Student Satisfaction Questionnaire Comments at ITB STIKOM Bali,” J. Appl. Inform. Comput., vol. 9, no. 3, pp. 794–802, Jun. 2025.
N. K. T. A. Saputri, I. G. A. Gunadi, and I. M. G. Sunarya, “Analisis Sentimen Pelayanan Daring di Fakultas Teknik dan Kejuruan Universitas Pendidikan Ganesha Menggunakan Algoritma Naïve Bayes dan LSTM: Sentiment Analysis of Online Services at the Engineering and Vocational Faculty of Ganesha Education University Using Naïve Bayes and LSTM Algorithms,” MALCOM Indones. J. Mach. Learn. Comput. Sci., vol. 4, no. 3, pp. 1120–1129, Jul. 2024, doi: 10.57152/malcom.v4i3.1336.
R. A. Raharjo, I. M. G. Sunarya, and D. G. H. Divayana, “Perbandingan Metode Naïve Bayes Classifier Dan Support Vector Machine Pada Kasus Analisis Sentimen Terhadap Data Vaksin Covid-19 Di Twitter,” Elkom J. Elektron. Dan Komput., vol. 15, no. 2, pp. 456–464, Dec. 2022, doi: 10.51903/elkom.v15i2.918.
E. Aditya, I. G. K. Astawa, K. G. Limbong, G. Indrawan, G. Indrawan, and M. A. O. Gunawan, “Analisis Sentimen Pengguna Sistem E-Kinerja Desa Kabupaten Jembrana Menggunakan Metode Naive Bayes,” J. Teknol. Dan Sist. Inf. Bisnis, vol. 7, no. 1, pp. 8–14, Jan. 2025, doi: 10.47233/jteksis.v7i1.1693.
N. M. A. J. Astari, Dewa Gede Hendra Divayana, and Gede Indrawan, “Analisis Sentimen Dokumen Twitter Mengenai Dampak Virus Corona Menggunakan Metode Naive Bayes Classifier,” J. Sist. Dan Inform. JSI, vol. 15, no. 1, pp. 27–29, Nov. 2020, doi: 10.30864/jsi.v15i1.332.
I. G. A. S. Sanjaya, I. M. Candiasa, and L. J. E. Dewi, “Analisis Sentimen Berbasis Aspek Kinerja Polri Menggunakan SVM dengan Pendekatan POS Tagging,” SINTECH Sci. Inf. Technol. J., vol. 7, no. 2, pp. 112–124, Aug. 2024, doi: 10.31598/sintechjournal.v7i2.1568.
G. Yunanda, D. Nurjanah, and S. Meliana, “Recommendation System from Microsoft News Data using TF-IDF and Cosine Similarity Methods,” Build. Inform. Technol. Sci. BITS, vol. 4, no. 1, Jun. 2022, doi: 10.47065/bits.v4i1.1670.
N. Kapoor, S. Vishal, and K. K. S., “Movie Recommendation System Using NLP Tools,” in 2020 5th International Conference on Communication and Electronics Systems (ICCES), Coimbatore, India: IEEE, Jun. 2020, pp. 883–888. doi: 10.1109/ICCES48766.2020.9137993.
D. Andreswari, D. Suranti, and D. A. Trianggara, “Application Of Text Mining In Grouping Thesis Topics Using TF-IDF Method Based On Thesis Abstract,” Dec. 2024.
M. A. Ariyanti, A. P. Wibawa, and U. Pujianto, “Metode term frequency - invers document frequency pada mekanisme pencarian judul skripsi,” TEKNO, vol. 28, no. 2, p. 177, Jul. 2019, doi: 10.17977/um034v28i2p177-190.
M. A. A. N. Fitroh, “Metode Term Frequency-Invers Document Frequency pada Mekanisme Pencarian Judul Skripsi Dalam Sistem Informasi Skripsi dan Tugas Akhir Jurusan Teknik Elektro Universitas Negeri Malang,” Skripsi, Universitas Negeri Malang, Malang, 2017.
D. M. Huda, “Rekomendasi Dokumen Skrispsi Terkait Menggunakan TFIDF pada Sistem Informasi Skripsi dan Tugas Akhir Jurusan Teknik Elektro Universitas Negeri Malang,” Skripsi, Universitas Negeri Malang, Malang, 2017.
D. A. U. Zaza, E. Umar, and F. E. O. Sanga, “Penerapan Text Mining Dalam Klasifikasi Judul Skripsi Mahasiswa Pada Universitas Stella Maris Sumba Menggunakan Metode Naïve Bayes,” vol. 02, no. 03, 2024.
C. Li, W. Li, Z. Tang, S. Li, and H. Xiang, “An Improved Term Weighting Method Based on Relevance Frequency for Text Classification,” Aug. 06, 2021. doi: 10.21203/rs.3.rs-680515/v1.
P. M. Mekontchou, A. Fotsoh, B. Batchakui, and E. Ella, “Information Retrieval in long documents: Word clustering approach for improving Semantics,” Feb. 20, 2023, arXiv: arXiv:2302.10150. Accessed: Jul. 10, 2024. [Online]. Available: http://arxiv.org/abs/2302.10150
Martius, Bahasa Indonesia Versi Mahasiswa Nonjurusan Bahasa Indonesia, 2nd ed. Pekanbaru - Riau: Asa Riau, 2018.
D. Marwah and J. Beel, “Term-Recency for TF-IDF, BM25 and USE Term Weighting”.
C. I. Muntean, F. M. Nardini, R. Perego, N. Tonellotto, and O. Frieder, “Weighting Passages Enhances Accuracy,” ACM Trans. Inf. Syst., vol. 39, no. 2, pp. 1–11, Apr. 2021, doi: 10.1145/3428687.
C. Kamphuis, A. P. De Vries, L. Boytsov, and J. Lin, “Which BM25 Do You Mean? A Large-Scale Reproducibility Study of Scoring Variants,” in Advances in Information Retrieval, vol. 12036, J. M. Jose, E. Yilmaz, J. Magalhães, P. Castells, N. Ferro, M. J. Silva, and F. Martins, Eds., in Lecture Notes in Computer Science, vol. 12036. , Cham: Springer International Publishing, 2020, pp. 28–34. doi: 10.1007/978-3-030-45442-5_4.
M. Li, D. N. Popa, J. Chagnon, Y. G. Cinar, and E. Gaussier, “The Power of Selecting Key Blocks with Local Pre-ranking for Long Document Information Retrieval,” ACM Trans. Inf. Syst., vol. 41, no. 3, pp. 1–35, Jul. 2023, doi: 10.1145/3568394.
R. Van Dinter, B. Tekinerdogan, and C. Catal, “Automation of systematic literature reviews: A systematic literature review,” Inf. Softw. Technol., vol. 136, p. 106589, Aug. 2021, doi: 10.1016/j.infsof.2021.106589.
D. Dessì, F. Osborne, D. R. Recupero, D. Buscaldi, and E. Motta, “Generating Knowledge Graphs by Employing Natural Language Processing and Machine Learning Techniques within the Scholarly Domain,” Future Gener. Comput. Syst., vol. 116, pp. 253–264, Mar. 2021, doi: 10.1016/j.future.2020.10.026.
H. Hairani, A. Anggrawan, A. I. Wathan, K. A. Latif, K. Marzuki, and M. Zulfikri, “The Abstract of Thesis Classifier by Using Naive Bayes Method,” in 2021 International Conference on Software Engineering & Computer Systems and 4th International Conference on Computational Science and Information Management (ICSECS-ICOCSIM), Pekan, Malaysia: IEEE, Aug. 2021, pp. 312–315. doi: 10.1109/ICSECS52883.2021.00063.
E. Melati, “COLLEGE STUDENT’S PROBLEMS IN WRITING PARAGRAPH ( A CASE STUDY AT FOURTH SEMESTER STUDENTS OF INFORMATICS MANAGEMENT OF AMIK MITRA GAMA),” ELP J. Engl. Lang. Pedagogy, vol. 5, no. 1, pp. 27–34, Jan. 2020, doi: 10.36665/elp.v5i1.245.
A. Hammache and M. Boughanem, “Term position‐based language model for information retrieval,” J. Assoc. Inf. Sci. Technol., vol. 72, no. 5, pp. 627–642, May 2021, doi: 10.1002/asi.24431.
K. Roberts et al., “Searching for scientific evidence in a pandemic: An overview of TREC-COVID,” J. Biomed. Inform., vol. 121, p. 103865, Sep. 2021, doi: 10.1016/j.jbi.2021.103865.
I. Soboroff, “Overview of TREC 2023,” National Institute of Standards and Technology, Gaithersburg, 500–342, May 2024. Accessed: Jul. 22, 2025. [Online]. Available: https://tsapps.nist.gov/publication/get_pdf.cfm?pub_id=957620
T. McDonnell, M. Kutlu, T. Elsayed, and M. Lease, “The Many Benefits of Annotator Rationales for Relevance Judgments,” in Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, Melbourne, Australia: International Joint Conferences on Artificial Intelligence Organization, Aug. 2017, pp. 4909–4913. doi: 10.24963/ijcai.2017/692.
M. L. McHugh, “Interrater reliability: the kappa statistic,” Biochem. Medica, pp. 276–282, 2012, doi: 10.11613/BM.2012.031.
P. A. A. Wijaya, I. M. Dika Anggara, I. K. Hendra Trinium Jaya, G. Indrawan, and I. M. Agus Oka Gunawan, “Comprehensive Analysis of Teacher Teaching Performance Through Sentiment and POS Tagging,” J. Ilm. Merpati Menara Penelit. Akad. Teknol. Inf., vol. 12, no. 3, p. 147, Dec. 2024, doi: 10.24843/JIM.2024.v12.i03.p02.
Moh. Heri Setiawan, I Gede Aris Gunadi, and Gede Indrawan, “Klasifikasi Pelayanan Kesehatan Berdasarkan Data Sentimen Pelayanan Kesehatan menggunakan Multiclass Support Vector Machine,” J. Sist. Dan Inform. JSI, vol. 17, no. 1, pp. 47–54, Apr. 2023, doi: 10.30864/jsi.v17i1.512.
J. Chen, C. Chen, and Y. Liang, “Optimized TF-IDF Algorithm with the Adaptive Weight of Position of Word,” in Proceedings of the 2016 2nd International Conference on Artificial Intelligence and Industrial Engineering (AIIE 2016), Beijing, China: Atlantis Press, 2016. doi: 10.2991/aiie-16.2016.28.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Putu Agung Ananta Wijaya, I Made Gede Sunarya, I Gede Aris Gunadi

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



