Hybrid Ensemble Retrieval-Augmented Generation for Indonesian Legal Consultation with Keyword Boosting

Hybrid Ensemble Retrieval-Augmented Generation for Indonesian Legal Consultation with Keyword Boosting

Authors

DOI:

https://doi.org/10.56741/jnest.v4i02.1042

Keywords:

Retrieval-Augmented Generation, RAG, Information Retrieval, FAISS, BM25, TF-IDF, Hybrid Ensemble Retrieval, Indonesian Legal QA System

Abstract

This study presents the design and evaluation of a fully local, hybrid ensemble Retrieval-Augmented Generation (RAG) system tailored for Indonesian legal consultation. By integrating sparse (BM25), dense (FAISS), and keyword-aware retrieval mechanisms, the system balances lexical, semantic, and domain-specific relevance to retrieve high-quality legal context. A curated dataset of 8,450 legal consultation articles was scraped from a trusted legal platform, cleaned through multi-stage pre-processing, and indexed for efficient retrieval. Retrieved documents are formatted into structured prompts and fed into locally hosted large language models (LLMs) using Ollama, allowing for complete offline operation. Experiments comparing five retrieval configurations TF-IDF, BM25, FAISS, ensemble BM25+FAISS, and ensemble with keyword boosting demonstrate that the hybrid ensemble with keyword boosting yields the most relevant and grounded answers. Both quantitative (retrieval score analysis) and qualitative (manual relevance rating) evaluations were performed, confirming the effectiveness of the ensemble strategy in improving answer quality. Additionally, the system achieves practical response times (12–20 seconds) on consumer-grade hardware without reliance on cloud services. This work makes a novel contribution by demonstrating that a hybrid ensemble retrieval framework, specifically tuned to the linguistic characteristics and retrieval challenges of Indonesian legal texts, can significantly enhance the performance of local RAG-based legal QA systems. Future directions include real-time indexing, fine-tuning of legal-domain LLMs, and extending the system to support other legal domains such as statutory law, regulations, and court rulings.

Downloads

Download data is not yet available.

Author Biographies

Suharyadi, Universitas Nusa Mandiri

Suharyadi is currently a Master's student in Computer Science at Universitas Nusa Mandiri, Jakarta. He has been working as an Information and Communication Technology Specialist at the Ministry of Finance, Republic of Indonesia, since 2010. His research interests include information retrieval, machine learning, deep learning, and technology innovation in public sector institutions. Correspondence: 14240001@nusamandiri.ac.id

Irwansyah Saputra, Universitas Nusa Mandiri

Irwansyah Saputra is a lecturer at the Faculty of Computer Science, Universitas Nusa Mandiri, Jakarta. He obtained his Doctoral degree from IPB University (Institut Pertanian Bogor), Indonesia, with a research focus in blockchain technology. His academic interests include distributed systems, blockchain applications, cybersecurity, and emerging technologies in computer science. He has published several works related to blockchain and continues to contribute to research and development in secure and decentralized computing systems. Correspondence: irwansyah.iys@nusamandiri.ac.id

References

P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” May 2020.

J. Martinez-Gil, “A survey on legal question–answering systems,” Comput Sci Rev, vol. 48, p. 100552, May 2023, doi: 10.1016/j.cosrev.2023.100552. DOI: https://doi.org/10.1016/j.cosrev.2023.100552

M. Douze et al., “The Faiss library,” Jan. 2024. DOI: https://doi.org/10.1109/TBDATA.2025.3618474

F. Liu, Z. Kang, and X. Han, “Optimizing RAG Techniques for Automotive Industry PDF Chatbots: A Case Study with Locally Deployed Ollama Models,” Aug. 2024. DOI: https://doi.org/10.1145/3707292.3707358

S. Robertson and H. Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Foundations and Trends® in Information Retrieval, vol. 3, no. 4, pp. 333–389, 2009, doi: 10.1561/1500000019. DOI: https://doi.org/10.1561/1500000019

I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, “LEGAL-BERT: The Muppets straight out of Law School,” in Findings of the Association for Computational Linguistics: EMNLP 2020, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 2898–2904. doi: 10.18653/v1/2020.findings-emnlp.261. DOI: https://doi.org/10.18653/v1/2020.findings-emnlp.261

F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Stroudsburg, PA, USA: International Committee on Computational Linguistics, 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66. DOI: https://doi.org/10.18653/v1/2020.coling-main.66

Z. Yang et al., “An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA,” Sep. 2021. DOI: https://doi.org/10.1609/aaai.v36i3.20215

J. Moreno Schneider et al., “Lynx: A knowledge-based AI service platform for content processing, enrichment and analysis for the legal domain,” Inf Syst, vol. 106, p. 101966, May 2022, doi: 10.1016/j.is.2021.101966. DOI: https://doi.org/10.1016/j.is.2021.101966

N. Xu, P. Wang, L. Chen, L. Pan, X. Wang, and J. Zhao, “Distinguish Confusing Law Articles for Legal Judgment Prediction,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 3086–3095. doi: 10.18653/v1/2020.acl-main.280. DOI: https://doi.org/10.18653/v1/2020.acl-main.280

V. Socatiyanurak et al., “LAW-U: Legal Guidance Through Artificial Intelligence Chatbot for Sexual Violence Victims and Survivors,” IEEE Access, vol. 9, pp. 131440–131461, 2021, doi: 10.1109/ACCESS.2021.3113172. DOI: https://doi.org/10.1109/ACCESS.2021.3113172

Hukumonline.com, “Kumpulan Tanya Jawab Atas Permasalahan Hukum | Hukumonline.” [Eng. Trans: Collections of Questions and Answers of Legal Issues]

X. Chen and S. Wiseman, “BM25 Query Augmentation Learned End-to-End,” May 2023.

H. Naveed et al., “A Comprehensive Overview of Large Language Models,” Jul. 2023.

T. Y. Zhuo et al., “Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models,” Jan. 2024.

A. R. Fabbri, W. Kryściński, B. McCann, C. Xiong, R. Socher, and D. Radev, “SummEval: Re-evaluating Summarization Evaluation,” Trans Assoc Comput Linguist, vol. 9, pp. 391–409, Apr. 2021, doi: 10.1162/tacl_a_00373. DOI: https://doi.org/10.1162/tacl_a_00373

C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries.”

Z. Dai and J. Callan, “Context-Aware Term Weighting For First Stage Passage Retrieval,” in Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA: ACM, Jul. 2020, pp. 1533–1536. doi: 10.1145/3397271.3401204. DOI: https://doi.org/10.1145/3397271.3401204

A. Abdallah, B. Piryani, and A. Jatowt, “Exploring the state of the art in legal QA systems,” J Big Data, vol. 10, no. 1, p. 127, Aug. 2023, doi: 10.1186/s40537-023-00802-8. DOI: https://doi.org/10.1186/s40537-023-00802-8

R. Campos, V. Mangaravite, A. Pasquali, A. Jorge, C. Nunes, and A. Jatowt, “YAKE! Keyword extraction from single documents using multiple local features,” Inf Sci (N Y), vol. 509, pp. 257–289, Jan. 2020, doi: 10.1016/j.ins.2019.09.013. DOI: https://doi.org/10.1016/j.ins.2019.09.013

B. Issa, M. B. Jasser, H. N. Chua, and M. Hamzah, “A Comparative Study on Embedding Models for Keyword Extraction Using KeyBERT Method,” in 2023 IEEE 13th International Conference on System Engineering and Technology (ICSET), IEEE, Oct. 2023, pp. 40–45. doi: 10.1109/ICSET59111.2023.10295108. DOI: https://doi.org/10.1109/ICSET59111.2023.10295108

Downloads

Published

2025-07-08

How to Cite

Suharyadi, & Saputra, I. (2025). Hybrid Ensemble Retrieval-Augmented Generation for Indonesian Legal Consultation with Keyword Boosting. Journal of Novel Engineering Science and Technology, 4(02), 71–85. https://doi.org/10.56741/jnest.v4i02.1042

Plaudit

Loading...