Multi-Stage Hybrid Retrieval with Neural Re-Ranking and Rule-Based Context Validation for Domain-Specific Retrieval-Augmented Generation

Multi-Stage Hybrid Retrieval with Neural Re-Ranking and Rule-Based Context Validation for Domain-Specific Retrieval-Augmented Generation

Authors

DOI:

https://doi.org/10.56741/jnest.v5i03.2329

Keywords:

BM25, Context Validation, Dense Retrieval, Hybrid Retrieval, Neural Re-Ranking, Retrieval-Augmented Generation, Question Answering

Abstract

Retrieval-Augmented Generation (RAG) has emerged as a widely adopted approach for grounding responses generated by large language models (LLMs) in external document corpora. However, conventional single-stage retrieval pipelines frequently fail to surface the most relevant context at top-ranked positions and lack explicit mechanisms to verify evidence adequacy prior to generation. This paper proposes a multi-stage RAG pipeline integrating three sequential components: hybrid retrieval combining BM25 and dense vector search via Reciprocal Rank Fusion (RRF), neural re-ranking using a cross-encoder transformer model, and a rule-based context validation gate that classifies retrieved context into three categories OK, WEAK, or ABSTAIN before answer generation. The system is evaluated on a domain-specific administrative corpus from the Indonesian AID Scholarship (TIAS) program using 50 annotated queries across four question types. Results demonstrate that hybrid retrieval improves recall at rank 10 (Recall@10) from 0.6391 to 0.8097 compared with BM25 alone, while neural re-ranking further increases mean reciprocal rank at rank 10 (MRR@10) from 0.6702 to 0.8040 (p = 0.0043, Wilcoxon signed-rank test). The complete pipeline achieves token-level F1 scores of 0.4429 and 0.4999 for the 4-billion-parameter (4B) and 12-billion-parameter (12B) model configurations, respectively, with a faithfulness score of 1.00. In comparison, the LLM-only baselines obtain token-level F1 scores below 0.09. The context validation mechanism correctly triggers ABSTAIN on 90% of out-of-domain queries, confirming its utility as a reliability control layer. These findings demonstrate that multi-stage retrieval with explicit context validation substantially improves both ranking quality and answer reliability in domain-specific RAG deployments.

Downloads

Download data is not yet available.

Author Biographies

Suharyadi, Universitas Nusa Mandiri

is currently a Master's student in Computer Science at Universitas Nusa Mandiri, Jakarta. He has been working as an Information and Communication Technology Specialist at the Ministry of Finance, Republic of Indonesia, since 2010. His research interests include information retrieval, machine learning, deep learning, and technology innovation in public sector institutions. (email: 14240001@nusamandiri.ac.id).

Irwansyah Saputra, Universitas Nusa Mandiri

is a lecturer at the Faculty of Computer Science, Universitas Nusa Mandiri, Jakarta. He obtained his Doctoral degree from IPB University (Institut Pertanian Bogor), Indonesia, with a research focus in blockchain technology. His academic interests include distributed systems, blockchain applications, cybersecurity, and emerging technologies in computer science. He has published several works related to blockchain and continues to contribute to research and development in secure and decentralized computing systems. (email: irwansyah.iys@nusamandiri.ac.id).

References

A. Vaswani et al., “Attention is all you need,” in 31st Conf. Adv. Neural Inf. Processing Systems, vol. 30, pp. 5998–6008, 2017. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf

Z. Ji et al., “Survey of hallucination in natural language generation,” ACM Comput. Surv., vol. 55, no. 12, pp. 1–38, Dec. 2023, doi: 10.1145/3571730.

P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in 34th Int. Conf. Adv. Neural Inf. Processing Systems, Curran Associates, Inc., pp. 9459–9474, 2020. [Online]. Available: https://dl.acm.org/doi/abs/10.5555/3495724.3496517

S. Robertson and H. Zaragoza, “The probabilistic relevance framework: BM25 and beyond,” Foundations and Trends® in Inf. Retrieval, vol. 3, no. 4, pp. 333–389, 2009, doi: 10.1561/1500000019.

V. Karpukhin et al., “Dense passage retrieval for open-domain question answering,” in Proc. the 2020 Conf. Empirical Methods in Natural Lang. Processing (EMNLP), pp. 6769–6781, 2020, doi: 10.18653/v1/2020.emnlp-main.550.

N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using siamese BERT-networks,” in Proc. of the 2019 Conf. on Empirical Methods in Natural Lang. Processing and the 9th Int. Joint Conf. on Natural Lang. Processing (EMNLP-IJCNLP), pp. 3982-3992, Aug. 2019. [Online]. Available: https://aclanthology.org/D19-1410/

Y. Gao et al., “Retrieval-augmented generation for large language models: A survey,” Comput. Lang. Mar. 2024, doi: 10.48550/arXiv.2312.10997.

J. Johnson, M. Douze, and H. Jegou, “Billion-scale similarity search with GPUs,” IEEE Trans. Big Data, vol. 7, no. 3, pp. 535–547, Jul. 2021, doi: 10.1109/TBDATA.2019.2921572.

M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, and P.-E. Mazaré, “The Faiss library,” IEEE Trans. on Big Data, vol. 12, no. 2, pp. 346-361, Oct. 2025, doi: 10.1109/TBDATA.2025.3618474.

L. Xiong et al., “Approximate nearest neighbor negative contrastive learning for dense text retrieval,” Inf. Retrieval, Oct. 2020, doi: 10.48550/arXiv.2007.00808.

G. Izacard et al., “Unsupervised dense information retrieval with contrastive learning,” Inf. Retrieval, Aug. 2022, doi: 10.48550/arXiv.2112.09118.

X. Ma, K. Sun, R. Pradeep, and J. Lin, “A replication study of dense passage retriever,” Comput. Lang., Apr. 2021, doi: 10.48550/arXiv.2104.05740.

G. V. Cormack, C. L. A. Clarke, and S. Buettcher, “Reciprocal rank fusion outperforms condorcet and individual rank learning methods,” in Proc. the 32nd Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, pp. 758–759, Jul. 2009, doi: 10.1145/1571941.1572114.

R. Nogueira and K. Cho, “Passage re-ranking with BERT,” Inf. Retrieval, Apr. 2020, doi: 10.48550/arXiv.1901.04085.

R. Nogueira, W. Yang, K. Cho, and J. Lin, “Multi-stage document ranking with BERT,” Oct. 2019, doi: 10.48550/arXiv.1910.14424.

O. Khattab and M. Zaharia, “ColBERT,” in Proc. 43rd Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, pp. 39–48, Jul. 2020, doi: 10.1145/3397271.3401075.

W. Sun et al., “Is ChatGPT good at search? Investigating large language models as re-ranking agents,” in Proc. 2023 Conf. Empirical Methods in Natural Lang. Processing, pp. 14918–14937, 2023, doi: 10.18653/v1/2023.emnlp-main.923.

G. Izacard and E. Grave, “Leveraging passage retrieval with generative models for open domain question answering,” in Proc. 16th Conf. Eur. Chapter Assoc. Comput. Linguist., pp. 874–880, 2021, doi: 10.18653/v1/2021.eacl-main.74.

O. Ram et al., “In-context retrieval-augmented language models,” Trans. Assoc. Comput. Linguist., vol. 11, pp. 1316–1331, Nov. 2023, doi: 10.1162/tacl_a_00605.

S. Siriwardhana, R. Weerasekera, E. Wen, T. Kaluarachchi, R. Rana, and S. Nanayakkara, “Improving the domain adaptation of Retrieval Augmented Generation (RAG) models for open domain question answering,” Trans. Assoc. Comput. Linguist., vol. 11, pp. 1–17, Jan. 2023, doi: 10.1162/tacl_a_00530.

Suharyadi and I. Saputra, “Hybrid ensemble retrieval-augmented generation for Indonesian legal consultation with keyword boosting,” J. Novel Eng. Sci. Technol., vol. 4, no. 02, pp. 71–85, Jul. 2025, doi: 10.56741/jnest.v4i02.1042.

O. Yoran, T. Wolfson, O. Ram, and J. Berant, “Making retrieval-augmented language models robust to irrelevant context,” Comput. Lang., May 2024, doi: 10.48550/arXiv.2310.01558.

W. Yu et al., “Chain-of-note: Enhancing robustness in retrieval-augmented language models,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 14672–14685, 2024, doi: 10.18653/v1/2024.emnlp-main.813.

A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-RAG: Learning to retrieve, generate, and critique through self-reflection,” Comput. Lang., Oct. 2023, doi: 10.48550/arXiv.2310.11511.

S.-Q. Yan, J.-C. Gu, Y. Zhu, and Z.-H. Ling, “Corrective retrieval augmented generation,” Comput. Lang., Oct. 2024, doi: 10.48550/arXiv.2401.15884.

Y. Wang, R. Ren, J. Li, W. X. Zhao, J. Liu, and J.-R. Wen, “REAR: A relevance-aware retrieval-augmented framework for open-domain question answering,” Nov. 2024. [Online]. Available: https://aclanthology.org/2024.emnlp-main.321/

D. Edge et al., “From local to global: A graph RAG approach to query-focused summarization,” Comput. Lang., Feb. 2025, doi: 10.48550/arXiv.2404.16130.

S. Jeong, J. Baek, S. Cho, S. J. Hwang, and J. Park, “Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity,” in Proc. 2024 Conf. North Am. Chapter Assoc. Comput. Linguist., pp. 7036–7050, 2024, doi: 10.18653/v1/2024.naacl-long.389.

D. Lee, Y. Jo, H. Park, and M. Lee, “Shifting from ranking to set selection for retrieval augmented generation,” in Proc. 63rd Annual Meeting Assoc. Comput. Linguist., pp. 17606–17619, 2025, doi: 10.18653/v1/2025.acl-long.861.

C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge University Press, 2009.

R. Baeza-Yates and B. Ribeiro-Neto, Modern Information Retrieval: The Concepts and Technology Behind Search. New York: ACM Press, 2011.

R. F. Woolson, “Wilcoxon signed‐rank test,” in Wiley Encyclopedia of Clinical Trials, Wiley, pp. 1–3, 2008, doi: 10.1002/9780471462422.eoct979.

I. T. Sorodoc, L. F. R. Ribeiro, R. Blloshmi, C. Davis, and A. de Gispert, “GaRAGe: A benchmark with grounding annotations for RAG evaluation,” in Findings of the Association for Computational Linguistics, pp. 17030–17049, 2025, doi: 10.18653/v1/2025.findings-acl.875.

Downloads

Published

2026-09-01

How to Cite

Suharyadi, & Saputra, I. (2026). Multi-Stage Hybrid Retrieval with Neural Re-Ranking and Rule-Based Context Validation for Domain-Specific Retrieval-Augmented Generation. Journal of Novel Engineering Science and Technology, 5(03), 174–189. https://doi.org/10.56741/jnest.v5i03.2329

Plaudit

Loading...