Retrieval-Augmented Generation and DeepSeek-R1 Large Language Models for Implementing Financial Report Chatbot for Companies Listed on the Indonesia Stock Exchange

Authors

  • Ivana Lucia Kharisma Informatics Engineering Study Program, Nusa Putra University, Indonesia
  • Haldies Gerhardien Pasya Informatics Engineering Study Program, Nusa Putra University, Indonesia
  • Alun Sujjada Informatics Engineering Study Program, Nusa Putra University, Indonesia

DOI:

https://doi.org/10.52436/1.jutif.2026.7.4.5562

Keywords:

Internet of Things, Object Detection, Raspberry Pi, YOLOv8, YOLOv8 Medium

Abstract

The Indonesia Stock Exchange (IDX) requires all listed companies to publish financial statements regularly to ensure transparency for investors and stakeholders. However, the increasing volume, complexity, and technical language of these reports create significant barriers for timely analysis, particularly for retail investors, regulators, and academic researchers. This limitation highlights the urgent need for intelligent systems capable of automating the extraction and interpretation of key financial information. This study aims to design and evaluate a Retrieval Augmented Generation (RAG) model to process financial reports, specifically the audited 2024 financial statements of Bank BCA and Bank Mandiri, in order to develop more efficient AI-based tools for managing complex financial documents. The research adopts a quantitative experimental approach using income statements from PT Bank Central Asia Tbk and PT Bank Mandiri (Persero) Tbk for the years 2023–2024. The methodology involves multimodal text extraction with Gemini Flash 2.0, preprocessing and cleaning using Regular Expressions (Regex), document chunking, and embedding generation with the multilingual-e5-small model. A Qdrant vector database is used for storage and retrieval, while DeepSeek-R1 serves as the transformer-based LLM for generating responses. Model performance was evaluated using BERTScore and ROUGE, supported by expert assessment. The system produced strong results, with BERTScore averages of 0.8074 (precision), 0.8253 (recall), and 0.8156 (F1-score). ROUGE evaluation yielded average scores of ROUGE-1: 0.5986 and ROUGE-2: 0.4531. Expert evaluation confirmed high parsing accuracy (94.6%) and the system’s ability to generate accurate, contextually relevant, and multilingual responses. Beyond its practical application, this study contributes to the advancement of scientific knowledge by providing an integrated framework for multimodal financial document processing using RAG-based LLMs. Overall, the study successfully developed a RAG-based LLM chatbot capable of effectively extracting, processing, and answering queries related to financial reports, demonstrating strong semantic alignment and supporting the development of more accessible AI-driven financial analysis tools.

Downloads

Download data is not yet available.

References

G. Josua Tulus Hartarto Program Studi Magister Kenotariatan Fakultas Hukum Universitas Airlangga Jl Dharmawangsa Dalam Selatan and K. Surabaya, “STATUS YURIDIS BURSA EFEK SEBAGAI PENGATUR KEGIATAN PERDAGANGAN PASAR MODAL,” Jilid, vol. 50, no. 2, pp. 143–150, 2021, doi: https://doi.org/10.14710/mmh.50.2.2021.143-150.

Albertus Daeli, Ribka Apriani Hutauruk, M.Budi Rifai, and Karina Silaen, “Analisis Laporan Keuangan Sebagai Penilai Kinerja Manajemen,” PPIMAN Pusat Publikasi Ilmu Manajemen, vol. 2, no. 3, pp. 158–168, Jul. 2024, doi: 10.59603/ppiman.v2i3.445.

Riani Tanjung and Rachel Erlita Anggrey Rani Sihite, “ANALISIS IMPLEMENTASI PSAK NO.1 PADA LAPORAN KEUANGAN PT. ANGKASA PURA II,” Jurnal Akuntansi, 2024, doi: 10.58457/akuntansi.v19i01.3753.

P. Rejison, “Identifying the importance of Financial Statements in Strategic Decision Making.”

L. Mohammed and M. Al-Juboori, “FINANCIAL STATEMENTS, THEIR COMPONENTS, OBJECTIVES, AND EFFECTIVENESS IN THE INSTITUTIONAL SYSTEM”, [Online]. Available: https://www.scholarexpress.net

T. Widia Nurdiani, M. Anas, and I. Sulistiana, “The Impact of Data Volume and Analytical Complexity in Big Data Technology on Financial Performance Prediction in Financial Companies in Indonesia Article Info ABSTRACT,” The Es Accounting and Finance, vol. 2, no. 01, pp. 64–76, doi: 10.58812/esaf.v2i01.

S. Cohen, F. Manes Rossi, X. Mamakou, and I. Brusca, “Financial accounting information presented with infographics: does it improve financial reporting understandability?,” Journal of Public Budgeting, Accounting and Financial Management, vol. 34, no. 6, pp. 263–295, 2022, doi: 10.1108/JPBAFM-11-2021-0163.

H. Ahaggach, L. Abrouk, and E. Lebon, “Information extraction from automotive reports for ontology population.”

M. Sheikh and S. Conlon, “A RULE-BASED SYSTEM TO EXTRACT FINANCIAL INFORMATION,” 2012. [Online]. Available: www.calcaria.net

Yi Xu, “Financial Statement Text Information Mining and Key Information Extraction Model Construction,” 2024. doi: 10.52783/jes.2743.

C.-C. Chen, H.-H. Huang, and H.-H. Chen, “NLP in FinTech Applications: Past, Present and Future.” [Online]. Available: www.hongkong-fintech.hk/;http://www.

A. Gupta, V. Dengre, H. A. Kheruwala, and M. Shah, “Comprehensive review of text-mining applications in finance,” Dec. 01, 2020, Springer Science and Business Media Deutschland GmbH. doi: 10.1186/s40854-020-00205-1.

L. Yang, Y. Li, J. Wang, and R. S. Sherratt, “Sentiment Analysis for E-Commerce Product Reviews in Chinese Based on Sentiment Lexicon and Deep Learning,” IEEE Access, vol. 8, pp. 23522–23530, 2020, doi: 10.1109/ACCESS.2020.2969854.

G. H. Soong and C. C. Tan, “Sentiment Analysis on 10-K Financial Reports using Machine Learning Approaches,” in IEEE 11th International Conference on System Engineering and Technology (ICSET), Shah Alam, Malaysia: IEEE, 2021, pp. 124–129.

N. Zhong and J. B. Ren, “Using sentiment analysis to study the relationship between subjective expression in financial reports and company performance,” Front. Psychol., vol. 13, Jul. 2022, doi: 10.3389/fpsyg.2022.949881.

Y. Ma, R. Mao, Q. Lin, P. Wu, and E. Cambria, “Multi-source Aggregated Classification for Stock Price Movement Prediction,” 2022.

Yu Ma, Rui Mao, Qika Lin, Peng Wu, and Erik Cambria, “Quantitative stock portfolio optimization by multi-task learning risk and return,” Information Fusion, vol. 104, 2024.

J. L. Bybee, “The Ghost in the Machine: Generating Beliefs with Large Language Models *.”

G.-Y. Choi et al., “Firm-Level Tax Audits: A Generative AI-Based Measurement.” [Online]. Available: http://ssrn.com/abstract=4645865Electroniccopyavailableat:https://ssrn.com/abstract=4645865

J. Rudolph, S. Tan, and S. Tan, “War of the chatbots: Bard, Bing Chat, ChatGPT, Ernie and beyond. The new AI gold rush and its impact on higher education,” Journal of Applied Learning and Teaching, vol. 6, no. 1, pp. 364–389, Jan. 2023, doi: 10.37074/jalt.2023.6.1.23.

V. K. and L. D. T. and P.-N. C. Nguyen Duy Dang Khoa and Mach, “Integrating Information Retrieval and LLMs: A Document Retrieval Chatbot in Education Settings,” in Intelligent Systems and Data Science, T.-N. and B. S. Thai-Nghe Nguyen and Do, Ed., Singapore: Springer Nature Singapore, 2026, pp. 461–475.

J. Genesis, “Retrieval-Augmented Text Generation: Methods, Challenges, and Applications,” Apr. 08, 2025. doi: 10.20944/preprints202504.0443.v1.

P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Apr. 2021, doi: 10.48550/arXiv.2005.11401.

S. Minaee et al., “Large Language Models: A Survey,” Mar. 2025, doi: 10.48550/arXiv.2402.06196.

Y. Gao et al., “Retrieval-Augmented Generation for Large Language Models: A Survey,” Mar. 2024, doi: 10.48550/arXiv.2312.10997.

Y. K. Chia et al., “M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework,” Nov. 2024, doi: 10.48550/arXiv.2411.06176.

X. Peng et al., “MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application,” Oct. 2025, doi: 10.48550/arXiv.2506.14028.

Z. Chen et al., “FinQA: A Dataset of Numerical Reasoning over Financial Data,” May 2022, doi: 10.48550/arXiv.2109.00122.

B. Saha, U. Saha, and M. Z. Malik, “QuIM-RAG: Advancing Retrieval-Augmented Generation With Inverted Question Matching for Enhanced QA Performance,” IEEE Access, vol. 12, pp. 185401–185410, 2024, doi: 10.1109/ACCESS.2024.3513155.

E. Oro, F. M. Granata, and M. Ruffolo, “A Comprehensive Evaluation of Embedding Models and LLMs for IR and QA Across English and Italian,” Big Data and Cognitive Computing, vol. 9, no. 5, May 2025, doi: 10.3390/bdcc9050141.

L. Pawlik, “LLM Selection and Vector Database Tuning: A Methodology for Enhancing RAG Systems,” Applied Sciences (Switzerland), vol. 15, no. 20, Oct. 2025, doi: 10.3390/app152010886.

J. Swacha and M. Gracel, “Retrieval-Augmented Generation (RAG) Chatbots for Education: A Survey of Applications,” Apr. 01, 2025, Multidisciplinary Digital Publishing Institute (MDPI). doi: 10.3390/app15084234.

R. M. Amodia, C. Tirnauca, and M. Zorrilla, “Original software publication RAGBOT CLI: a Python library for running and evaluating retrieval-augmented generation chatbots,” 2025, doi: 10.5281/zenodo.17300635.

T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “BERTScore: Evaluating Text Generation with BERT,” Feb. 2020, [Online]. Available: http://arxiv.org/abs/1904.09675

C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries.”

A. R. Fabbri, W. Kryściński, B. McCann, C. Xiong, R. Socher, and D. Radev, “SummEval: Re-evaluating Summarization Evaluation,” Feb. 2021, [Online]. Available: http://arxiv.org/abs/2007.12626

G. A. N. Lubis, R. Purnamasari, and K. Saleh, “Designing a Real-Time TOXMAP Backend Based on FastAPI and Firebase for B3 Waste,” bit-Tech, vol. 8, no. 1, pp. 1159–1167, Aug. 2025, doi: 10.32877/bt.v8i1.2888.

M. Rokita, M. Modrzejewski, and P. Rokita, “PyBrook—A Python framework for processing and visualising real-time data,” SoftwareX, vol. 30, May 2025, doi: 10.1016/j.softx.2025.102116.

Praveen Borra, “A Survey of Google Cloud Platform (GCP): Features, Services, and Applications,” International Journal of Advanced Research in Science, Communication and Technology, pp. 191–199, Jun. 2024, doi: 10.48175/ijarsct-18922.

Jaakko Soitinaho, “Migrating Enterprise Single-Page Application to Server-side Renderable Application,” Master Programme in Computer, 2024.

G. S. Rao, T. N. Lakshmi, R. S. Parvathi, S. Chandra Sekhar, and P. L. Rao, “314 International Journal for Modern Trends in Science and Technology Automated Image Enhancement for Historical Document preservation and Analysis with Gemini AI,” Analysis with Gemini AI. International Journal for Modern Trends in Science and Technology, vol. 11, no. 04, pp. 314–321, doi: 10.5281/zenodo.15131020.

Q. Chen, X. Wang, X. Ye, G. Durrett, and I. Dillig, “Multi-modal synthesis of regular expressions,” in Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), Association for Computing Machinery, Jun. 2020, pp. 487–502. doi: 10.1145/3385412.3385988.

V. Bulgakov and A. Segal, “Dimensionality Reduction in Sentence Transformer Vector Databases with Fast Fourier Transform,” Apr. 2024, doi: 10.48550/arXiv.2404.06278.

B. Plale, S. N. Jyesta, and S. Withana, “Vector embedding of multi-modal texts: a tool for discovery?,” Sep. 2025, doi: 10.48550/arXiv.2509.08216.

M. Yonathan, “Access-Controlled Semantic Search: Implementing Role-Based Filtering in Vector Databases for Enterprise Document Management.” doi: doi.org/10.5281/zenodo.15717386.

G. Izacard and E. Grave, “Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering.”

J. Choe, J. Kim, and W. Jung, “Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents,” Nov. 2025, doi: 10.18653/v1/2025.findings-acl.855.

Y. Zhao et al., “Optimizing LLM Based Retrieval Augmented Generation Pipelines in the Financial Domain.” [Online]. Available: https://platform.openai.com/docs/guides/embeddings

Additional Files

Published

2026-08-17

How to Cite

[1]
I. Lucia Kharisma, H. Gerhardien Pasya, and A. Sujjada, “Retrieval-Augmented Generation and DeepSeek-R1 Large Language Models for Implementing Financial Report Chatbot for Companies Listed on the Indonesia Stock Exchange”, J. Tek. Inform. (JUTIF), vol. 7, no. 4, pp. 3625–3640, Aug. 2026.