Hybrid Ensemble Retrieval-Augmented Generation for Indonesian Legal Consultation with Keyword Boosting
DOI:
https://doi.org/10.56741/jnest.v4i02.1042Keywords:
Retrieval-Augmented Generation, RAG, Information Retrieval, FAISS, BM25, TF-IDF, Hybrid Ensemble Retrieval, Indonesian Legal QA SystemAbstract
This study presents the design and evaluation of a fully local, hybrid ensemble Retrieval-Augmented Generation (RAG) system tailored for Indonesian legal consultation. By integrating sparse (BM25), dense (FAISS), and keyword-aware retrieval mechanisms, the system balances lexical, semantic, and domain-specific relevance to retrieve high-quality legal context. A curated dataset of 8,450 legal consultation articles was scraped from a trusted legal platform, cleaned through multi-stage pre-processing, and indexed for efficient retrieval. Retrieved documents are formatted into structured prompts and fed into locally hosted large language models (LLMs) using Ollama, allowing for complete offline operation. Experiments comparing five retrieval configurations TF-IDF, BM25, FAISS, ensemble BM25+FAISS, and ensemble with keyword boosting demonstrate that the hybrid ensemble with keyword boosting yields the most relevant and grounded answers. Both quantitative (retrieval score analysis) and qualitative (manual relevance rating) evaluations were performed, confirming the effectiveness of the ensemble strategy in improving answer quality. Additionally, the system achieves practical response times (12–20 seconds) on consumer-grade hardware without reliance on cloud services. This work makes a novel contribution by demonstrating that a hybrid ensemble retrieval framework, specifically tuned to the linguistic characteristics and retrieval challenges of Indonesian legal texts, can significantly enhance the performance of local RAG-based legal QA systems. Future directions include real-time indexing, fine-tuning of legal-domain LLMs, and extending the system to support other legal domains such as statutory law, regulations, and court rulings.
Downloads
References
P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” May 2020.
J. Martinez-Gil, “A survey on legal question–answering systems,” Comput Sci Rev, vol. 48, p. 100552, May 2023, doi: 10.1016/j.cosrev.2023.100552. DOI: https://doi.org/10.1016/j.cosrev.2023.100552
M. Douze et al., “The Faiss library,” Jan. 2024. DOI: https://doi.org/10.1109/TBDATA.2025.3618474
F. Liu, Z. Kang, and X. Han, “Optimizing RAG Techniques for Automotive Industry PDF Chatbots: A Case Study with Locally Deployed Ollama Models,” Aug. 2024. DOI: https://doi.org/10.1145/3707292.3707358
S. Robertson and H. Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Foundations and Trends® in Information Retrieval, vol. 3, no. 4, pp. 333–389, 2009, doi: 10.1561/1500000019. DOI: https://doi.org/10.1561/1500000019
I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, “LEGAL-BERT: The Muppets straight out of Law School,” in Findings of the Association for Computational Linguistics: EMNLP 2020, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 2898–2904. doi: 10.18653/v1/2020.findings-emnlp.261. DOI: https://doi.org/10.18653/v1/2020.findings-emnlp.261
F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Stroudsburg, PA, USA: International Committee on Computational Linguistics, 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66. DOI: https://doi.org/10.18653/v1/2020.coling-main.66
Z. Yang et al., “An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA,” Sep. 2021. DOI: https://doi.org/10.1609/aaai.v36i3.20215
J. Moreno Schneider et al., “Lynx: A knowledge-based AI service platform for content processing, enrichment and analysis for the legal domain,” Inf Syst, vol. 106, p. 101966, May 2022, doi: 10.1016/j.is.2021.101966. DOI: https://doi.org/10.1016/j.is.2021.101966
N. Xu, P. Wang, L. Chen, L. Pan, X. Wang, and J. Zhao, “Distinguish Confusing Law Articles for Legal Judgment Prediction,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 3086–3095. doi: 10.18653/v1/2020.acl-main.280. DOI: https://doi.org/10.18653/v1/2020.acl-main.280
V. Socatiyanurak et al., “LAW-U: Legal Guidance Through Artificial Intelligence Chatbot for Sexual Violence Victims and Survivors,” IEEE Access, vol. 9, pp. 131440–131461, 2021, doi: 10.1109/ACCESS.2021.3113172. DOI: https://doi.org/10.1109/ACCESS.2021.3113172
Hukumonline.com, “Kumpulan Tanya Jawab Atas Permasalahan Hukum | Hukumonline.” [Eng. Trans: Collections of Questions and Answers of Legal Issues]
X. Chen and S. Wiseman, “BM25 Query Augmentation Learned End-to-End,” May 2023.
H. Naveed et al., “A Comprehensive Overview of Large Language Models,” Jul. 2023.
T. Y. Zhuo et al., “Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models,” Jan. 2024.
A. R. Fabbri, W. Kryściński, B. McCann, C. Xiong, R. Socher, and D. Radev, “SummEval: Re-evaluating Summarization Evaluation,” Trans Assoc Comput Linguist, vol. 9, pp. 391–409, Apr. 2021, doi: 10.1162/tacl_a_00373. DOI: https://doi.org/10.1162/tacl_a_00373
C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries.”
Z. Dai and J. Callan, “Context-Aware Term Weighting For First Stage Passage Retrieval,” in Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA: ACM, Jul. 2020, pp. 1533–1536. doi: 10.1145/3397271.3401204. DOI: https://doi.org/10.1145/3397271.3401204
A. Abdallah, B. Piryani, and A. Jatowt, “Exploring the state of the art in legal QA systems,” J Big Data, vol. 10, no. 1, p. 127, Aug. 2023, doi: 10.1186/s40537-023-00802-8. DOI: https://doi.org/10.1186/s40537-023-00802-8
R. Campos, V. Mangaravite, A. Pasquali, A. Jorge, C. Nunes, and A. Jatowt, “YAKE! Keyword extraction from single documents using multiple local features,” Inf Sci (N Y), vol. 509, pp. 257–289, Jan. 2020, doi: 10.1016/j.ins.2019.09.013. DOI: https://doi.org/10.1016/j.ins.2019.09.013
B. Issa, M. B. Jasser, H. N. Chua, and M. Hamzah, “A Comparative Study on Embedding Models for Keyword Extraction Using KeyBERT Method,” in 2023 IEEE 13th International Conference on System Engineering and Technology (ICSET), IEEE, Oct. 2023, pp. 40–45. doi: 10.1109/ICSET59111.2023.10295108. DOI: https://doi.org/10.1109/ICSET59111.2023.10295108
Downloads
Published
How to Cite
Issue
Section
Categories
License
Copyright (c) 2025 Journal of Novel Engineering Science and Technology

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.






















