Multi-Stage Hybrid Retrieval with Neural Re-Ranking and Rule-Based Context Validation for Domain-Specific Retrieval-Augmented Generation
DOI:
https://doi.org/10.56741/jnest.v5i03.2329Keywords:
BM25, Context Validation, Dense Retrieval, Hybrid Retrieval, Neural Re-Ranking, Retrieval-Augmented Generation, Question AnsweringAbstract
Retrieval-Augmented Generation (RAG) has emerged as a widely adopted approach for grounding responses generated by large language models (LLMs) in external document corpora. However, conventional single-stage retrieval pipelines frequently fail to surface the most relevant context at top-ranked positions and lack explicit mechanisms to verify evidence adequacy prior to generation. This paper proposes a multi-stage RAG pipeline integrating three sequential components: hybrid retrieval combining BM25 and dense vector search via Reciprocal Rank Fusion (RRF), neural re-ranking using a cross-encoder transformer model, and a rule-based context validation gate that classifies retrieved context into three categories OK, WEAK, or ABSTAIN before answer generation. The system is evaluated on a domain-specific administrative corpus from the Indonesian AID Scholarship (TIAS) program using 50 annotated queries across four question types. Results demonstrate that hybrid retrieval improves recall at rank 10 (Recall@10) from 0.6391 to 0.8097 compared with BM25 alone, while neural re-ranking further increases mean reciprocal rank at rank 10 (MRR@10) from 0.6702 to 0.8040 (p = 0.0043, Wilcoxon signed-rank test). The complete pipeline achieves token-level F1 scores of 0.4429 and 0.4999 for the 4-billion-parameter (4B) and 12-billion-parameter (12B) model configurations, respectively, with a faithfulness score of 1.00. In comparison, the LLM-only baselines obtain token-level F1 scores below 0.09. The context validation mechanism correctly triggers ABSTAIN on 90% of out-of-domain queries, confirming its utility as a reliability control layer. These findings demonstrate that multi-stage retrieval with explicit context validation substantially improves both ranking quality and answer reliability in domain-specific RAG deployments.
Downloads
References
A. Vaswani et al., “Attention is all you need,” in 31st Conf. Adv. Neural Inf. Processing Systems, vol. 30, pp. 5998–6008, 2017. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
Z. Ji et al., “Survey of hallucination in natural language generation,” ACM Comput. Surv., vol. 55, no. 12, pp. 1–38, Dec. 2023, doi: 10.1145/3571730.
P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in 34th Int. Conf. Adv. Neural Inf. Processing Systems, Curran Associates, Inc., pp. 9459–9474, 2020. [Online]. Available: https://dl.acm.org/doi/abs/10.5555/3495724.3496517
S. Robertson and H. Zaragoza, “The probabilistic relevance framework: BM25 and beyond,” Foundations and Trends® in Inf. Retrieval, vol. 3, no. 4, pp. 333–389, 2009, doi: 10.1561/1500000019.
V. Karpukhin et al., “Dense passage retrieval for open-domain question answering,” in Proc. the 2020 Conf. Empirical Methods in Natural Lang. Processing (EMNLP), pp. 6769–6781, 2020, doi: 10.18653/v1/2020.emnlp-main.550.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using siamese BERT-networks,” in Proc. of the 2019 Conf. on Empirical Methods in Natural Lang. Processing and the 9th Int. Joint Conf. on Natural Lang. Processing (EMNLP-IJCNLP), pp. 3982-3992, Aug. 2019. [Online]. Available: https://aclanthology.org/D19-1410/
Y. Gao et al., “Retrieval-augmented generation for large language models: A survey,” Comput. Lang. Mar. 2024, doi: 10.48550/arXiv.2312.10997.
J. Johnson, M. Douze, and H. Jegou, “Billion-scale similarity search with GPUs,” IEEE Trans. Big Data, vol. 7, no. 3, pp. 535–547, Jul. 2021, doi: 10.1109/TBDATA.2019.2921572.
M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, and P.-E. Mazaré, “The Faiss library,” IEEE Trans. on Big Data, vol. 12, no. 2, pp. 346-361, Oct. 2025, doi: 10.1109/TBDATA.2025.3618474.
L. Xiong et al., “Approximate nearest neighbor negative contrastive learning for dense text retrieval,” Inf. Retrieval, Oct. 2020, doi: 10.48550/arXiv.2007.00808.
G. Izacard et al., “Unsupervised dense information retrieval with contrastive learning,” Inf. Retrieval, Aug. 2022, doi: 10.48550/arXiv.2112.09118.
X. Ma, K. Sun, R. Pradeep, and J. Lin, “A replication study of dense passage retriever,” Comput. Lang., Apr. 2021, doi: 10.48550/arXiv.2104.05740.
G. V. Cormack, C. L. A. Clarke, and S. Buettcher, “Reciprocal rank fusion outperforms condorcet and individual rank learning methods,” in Proc. the 32nd Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, pp. 758–759, Jul. 2009, doi: 10.1145/1571941.1572114.
R. Nogueira and K. Cho, “Passage re-ranking with BERT,” Inf. Retrieval, Apr. 2020, doi: 10.48550/arXiv.1901.04085.
R. Nogueira, W. Yang, K. Cho, and J. Lin, “Multi-stage document ranking with BERT,” Oct. 2019, doi: 10.48550/arXiv.1910.14424.
O. Khattab and M. Zaharia, “ColBERT,” in Proc. 43rd Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, pp. 39–48, Jul. 2020, doi: 10.1145/3397271.3401075.
W. Sun et al., “Is ChatGPT good at search? Investigating large language models as re-ranking agents,” in Proc. 2023 Conf. Empirical Methods in Natural Lang. Processing, pp. 14918–14937, 2023, doi: 10.18653/v1/2023.emnlp-main.923.
G. Izacard and E. Grave, “Leveraging passage retrieval with generative models for open domain question answering,” in Proc. 16th Conf. Eur. Chapter Assoc. Comput. Linguist., pp. 874–880, 2021, doi: 10.18653/v1/2021.eacl-main.74.
O. Ram et al., “In-context retrieval-augmented language models,” Trans. Assoc. Comput. Linguist., vol. 11, pp. 1316–1331, Nov. 2023, doi: 10.1162/tacl_a_00605.
S. Siriwardhana, R. Weerasekera, E. Wen, T. Kaluarachchi, R. Rana, and S. Nanayakkara, “Improving the domain adaptation of Retrieval Augmented Generation (RAG) models for open domain question answering,” Trans. Assoc. Comput. Linguist., vol. 11, pp. 1–17, Jan. 2023, doi: 10.1162/tacl_a_00530.
Suharyadi and I. Saputra, “Hybrid ensemble retrieval-augmented generation for Indonesian legal consultation with keyword boosting,” J. Novel Eng. Sci. Technol., vol. 4, no. 02, pp. 71–85, Jul. 2025, doi: 10.56741/jnest.v4i02.1042.
O. Yoran, T. Wolfson, O. Ram, and J. Berant, “Making retrieval-augmented language models robust to irrelevant context,” Comput. Lang., May 2024, doi: 10.48550/arXiv.2310.01558.
W. Yu et al., “Chain-of-note: Enhancing robustness in retrieval-augmented language models,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 14672–14685, 2024, doi: 10.18653/v1/2024.emnlp-main.813.
A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-RAG: Learning to retrieve, generate, and critique through self-reflection,” Comput. Lang., Oct. 2023, doi: 10.48550/arXiv.2310.11511.
S.-Q. Yan, J.-C. Gu, Y. Zhu, and Z.-H. Ling, “Corrective retrieval augmented generation,” Comput. Lang., Oct. 2024, doi: 10.48550/arXiv.2401.15884.
Y. Wang, R. Ren, J. Li, W. X. Zhao, J. Liu, and J.-R. Wen, “REAR: A relevance-aware retrieval-augmented framework for open-domain question answering,” Nov. 2024. [Online]. Available: https://aclanthology.org/2024.emnlp-main.321/
D. Edge et al., “From local to global: A graph RAG approach to query-focused summarization,” Comput. Lang., Feb. 2025, doi: 10.48550/arXiv.2404.16130.
S. Jeong, J. Baek, S. Cho, S. J. Hwang, and J. Park, “Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity,” in Proc. 2024 Conf. North Am. Chapter Assoc. Comput. Linguist., pp. 7036–7050, 2024, doi: 10.18653/v1/2024.naacl-long.389.
D. Lee, Y. Jo, H. Park, and M. Lee, “Shifting from ranking to set selection for retrieval augmented generation,” in Proc. 63rd Annual Meeting Assoc. Comput. Linguist., pp. 17606–17619, 2025, doi: 10.18653/v1/2025.acl-long.861.
C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge University Press, 2009.
R. Baeza-Yates and B. Ribeiro-Neto, Modern Information Retrieval: The Concepts and Technology Behind Search. New York: ACM Press, 2011.
R. F. Woolson, “Wilcoxon signed‐rank test,” in Wiley Encyclopedia of Clinical Trials, Wiley, pp. 1–3, 2008, doi: 10.1002/9780471462422.eoct979.
I. T. Sorodoc, L. F. R. Ribeiro, R. Blloshmi, C. Davis, and A. de Gispert, “GaRAGe: A benchmark with grounding annotations for RAG evaluation,” in Findings of the Association for Computational Linguistics, pp. 17030–17049, 2025, doi: 10.18653/v1/2025.findings-acl.875.
Downloads
Published
How to Cite
Issue
Section
Categories
License
Copyright (c) 2026 Journal of Novel Engineering Science and Technology

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.






















