← All RAG & retrieval lessons

RAG & retrieval / HANDS-ON LESSON

Combine two retrieval rankings

Use reciprocal rank fusion to merge keyword and semantic candidates without mixing their raw scores.

IntermediateAbout 20 min5 practice questionsReviewed 2026-10-05SaveMyToken editorial

WHAT YOU WILL BUILD

A deduplicated ranking with a visible, reproducible scoring rule.

A keyword engine and a semantic retriever can return different score scales. Reciprocal rank fusion uses positions instead: each list contributes 1 / (constant + rank), with rank starting at one. It rewards agreement without pretending the original scores are comparable.

Before you start

Use a text editor and Node.js 22 or newer ↗. Check your version with node --version. Save the downloaded file in an empty folder, then open a terminal in that folder.

All data is included. No packages, account, or API key are required. The two input rankings are supplied teaching data. The script computes their fusion; it does not measure a keyword engine or embedding model.

Download the lab (.mjs)

FOLLOW ALONG

Work through the example.

  1. 01

    Inspect the candidate IDs

    The first ranking favors sku-42; the second favors returns. Notice that returns and sku-42 occur in both lists. A document must keep the same ID across retrieval systems.

  2. 02

    Run the fusion

    Run the downloaded file. returns comes first, followed by sku-42. Print scores to inspect the sum contributed by each source ranking.

  3. 03

    Change one ranking

    Swap returns and sku-42 in the second list. Predict the new top document before running. Update the expected ranking assertion after checking the arithmetic.

  4. 04

    Compare against a labeled question

    Decide which documents are relevant to a specific question, then compare each individual list and the fused list at the same K. A changed order is useful only if it retrieves better evidence.

THE COMPLETE LAB

Run it locally.

Run this command from the folder containing your downloaded file:

node rag-rank-fusion.mjs

View or copy the complete JavaScript
// SaveMyToken local lab. Run with Node.js 22 or newer.
import assert from "node:assert/strict";

// Two supplied rankings: no embedding model is being run here.
const lexical = ["sku-42", "returns", "shipping"];
const semantic = ["returns", "refunds", "sku-42"];
const rankConstant = 60;
const scores = new Map();
for (const list of [lexical, semantic]) {
  [...new Set(list)].forEach((id, index) => {
    scores.set(id, (scores.get(id) ?? 0) + 1 / (rankConstant + index + 1));
  });
}
const ranked = [...scores].sort((a, b) => b[1] - a[1] || a[0].localeCompare(b[0]));
const result = ranked.map(([id]) => id);
console.log("Fused ranking:", result.join(", "));
assert.deepEqual(result, ["returns", "sku-42", "refunds", "shipping"]);
assert.equal(new Set(result).size, result.length);

Expected output for the unchanged example

Fused ranking: returns, sku-42, refunds, shipping

The assertions also check the baseline behavior. After changing an input, predict the result and update the relevant assertion.

NOW CHANGE ONE THING

Make the example your own.

Add a third supplied ranking and compare rankConstant values 1 and 60. Record which results change; keep the same relevance labels throughout.

You are done when…

Every candidate appears once, ranks start at one, and you can explain the winning document's score.

If something goes wrong

Unexpected duplicates usually mean inconsistent document IDs. Ties here use an ID sort for reproducibility; the tie-break is not an extra relevance signal.

EXPLAIN WHAT YOU LEARNED

Interview practice

Try answering aloud before opening the reference answer. These are original learning questions, not a record of any employer's interviews.

01Why combine keyword and semantic retrieval?

They can find complementary evidence: exact identifiers matter for keyword matches, while meaning-based retrieval can help with paraphrases. Evaluate their combination against your own questions.

Watch for: Hybrid retrieval is not automatically better on every query.

02Why not add their raw scores?

The scales and distributions may differ, allowing one retriever to dominate. Use a calibrated combination or a rank-based method with explicit assumptions.

Watch for: A score of 0.8 from one system is not necessarily comparable to 0.8 from another.

03What does the RRF constant do?

It controls how much the highest positions dominate each contribution. A smaller constant makes early rank differences more influential.

Watch for: It is not the number of retrieved documents or a similarity threshold.

04How is reranking different from fusion?

Fusion merges candidate orderings. A reranker evaluates candidate relevance to the query with another scoring procedure. You may fuse first and rerank the resulting shortlist.

Watch for: A reranker cannot recover evidence absent from its candidate set.

05How do you avoid duplicate evidence?

Use stable document or passage IDs, deduplicate within each result list, then merge cross-list contributions. Also inspect near-duplicate passages before constructing the final context.

Watch for: Deduplicating unrelated passages only because their titles match can remove evidence.

Practice more rag & retrieval questions

Sources and further reading

The explanation and local exercises were written for SaveMyToken. These references support the underlying concepts; the sample outputs describe only the supplied examples.

Elasticsearch: reciprocal rank fusion ↗