A dense retriever misses an exact part number. A keyword retriever finds the number but misses a question phrased in different language. Combining them may help, but it also changes the candidate set and can add noise. “Hybrid” is a method to test, not a quality guarantee.
Start by defining what evidence should be retrieved for a task. Generation cannot reliably recover an answer from evidence that never enters the candidate set, and a reranker cannot promote a passage it never sees.
Measure coverage before answer style
Build relevance judgments at the document or passage level. For a question needing a general rule and an exception, mark both as necessary evidence. Report candidate recall at a stated cutoff, not merely whether one relevant document appeared somewhere.
BEIR evaluates retrieval across heterogeneous tasks and domains. It is useful evidence that retrieval performance depends on the setting. It does not make one ranking system universally best for an enterprise corpus with its own terminology and permissions.
Fuse rankings, not incompatible raw scores
Lexical and dense scores often have different scales. A simple rank-based method is reciprocal rank fusion, or RRF: add 1/(c+rank) for each ranking in which a document appears. The constant c dampens the influence of the top ranks; it is a configuration choice, not evidence of correctness.
For a synthetic example with c=60, document A appears at lexical rank 1 and dense rank 10, while document B appears at rank 2 in both. A scores 1/61+1/70≈0.03068. B scores 1/62+1/62≈0.03226, so B ranks above A. The calculation illustrates the rule. It does not show that B actually contains better evidence.
Keep the fusion configuration visible
Specify the number of candidates from each retriever, the fusion rule, any rank window, and tie handling. If the lexical list contains 10 documents and the dense list contains 100, the comparison is not symmetrical. Evaluate the configuration on representative queries before tuning it against a release holdout.
The Elasticsearch RRF documentation gives the formula and implementation behavior. Use its current API contract for implementation details; the formula alone does not define every retrieval setting.
Rerank the evidence, not just the wording
A reranker can improve the order of a candidate set, but it consumes work per candidate and may still prefer a plausible wrong passage. Score source scope, effective version, and evidence completeness alongside relevance where the task requires them.
Evaluate lexical-only, dense-only, fused, and fused-plus-reranked variants. Record the evidence lost at each stage, latency, and generation outcomes. A better ranking score is useful only if the answer or task improves under the same source contract.
Apply permission boundaries before exposure
Filter access before restricted material enters model context and avoid sharing cached candidates across users who have different permissions. For consequential tasks, apply authorization at the resource owner too. A permission filter in search is not the only security boundary.
Include deleted sources, revoked access, similar identifiers, multilingual wording, and empty queries in the test suite. Check that a rare exact match is not drowned out by generic semantic similarity.
Choose the smallest useful retrieval pipeline
If keyword search resolves the exact-record workload, keep it. If fusion improves paraphrased queries but hurts identifiers, use task-aware routing or a validated configuration rather than claiming one global win. Preserve an unanswerable path when evidence is absent.
The candidate set is the first place where retrieval can be honestly evaluated. Once that boundary is clear, generation becomes easier to diagnose and harder to credit for evidence it never had.