Stop Vibing Your RAG Pipeline: Recall@k, MRR, and NDCG in an Afternoon

Most teams that ship a retrieval-augmented generation feature never measure it. The demo looks good, a few hand-picked questions return convincing answers, and the system goes to production carrying an invisible error rate. Then a user asks something where the retrieval layer surfaces the wrong document — the model answers fluently and confidently from garbage context — and support tickets start arriving. The uncomfortable part is that the failure was measurable all along: not by vibes, but by a handful of retrieval metrics that take an afternoon to implement.

This post covers the evaluation stack for RAG systems in the order you should adopt it. First the retrieval metrics — recall@k, MRR, and NDCG — which need only a labeled question set and tell you whether the right documents reach the model at all. Then generation-side checks: groundedness (does the answer stick to the retrieved context?) and relevance (does it answer the question?). Finally, a working harness in Python that computes all of them against a small golden dataset, so the pipeline has numbers attached to it before the next prompt change ships.

Start With a Golden Dataset

Every metric below consumes the same input: a set of questions where you know, ahead of time, which documents should be retrieved and ideally what a good answer looks like. Fifty to two hundred questions is enough to catch regressions. Building it is tedious but cheap if you mine it from real traffic — pull actual user questions from logs, attach the documents that were retrieved, and have someone spend an hour labeling which retrievals were correct and which answers were acceptable. Synthetic question generation (have an LLM write questions whose answers are in specific documents) works as a bootstrap, but expect a distribution shift: real users phrase things worse than synthetic generators do.

golden: list[dict] = [
    {
        "question": "What is the maximum file upload size on the Pro plan?",
        "relevant_ids": {"docs/plans-0012"},   # chunk/document IDs that SHOULD be retrieved
        "expected_answer": "Pro plan uploads are capped at 5 GB per file.",
    },
    {
        "question": "How do I rotate my API key?",
        "relevant_ids": {"docs/auth-0003", "docs/auth-0007"},
    },
]

Note that some questions have several relevant documents. The retrieval metrics below handle that differently, which is exactly why you want more than one of them.

Recall@k: Did the Right Document Even Show Up?

Recall@k is the single most important RAG metric, because everything downstream is bounded by it. It asks: for the top k retrieved documents, what fraction of the relevant documents were included? If recall@10 is 0.8, no amount of prompt engineering recovers the 20% of questions where the evidence never reached the model.

def recall_at_k(retrieved: list[str], relevant: set[str], k: int) -> float:
    top_k = set(retrieved[:k])
    if not relevant:
        return 0.0
    return len(top_k & relevant) / len(relevant)

Two practical notes. First, choose k to match your actual pipeline: if you pass five chunks to the model, recall@5 is the number that matters, not recall@20. Second, recall says nothing about ranking among the retrieved set — a system that puts the gold document at position 9 of 10 and one that puts it first score identically. That is what the next two metrics fix.

MRR: How Fast Do Users Hit the Right Answer?

Mean Reciprocal Rank scores the position of the first relevant document: 1 divided by its rank, averaged over the question set. It models the common RAG situation where one document carries the answer and the rest are filler. A system that retrieves the right chunk first scores 1.0; if it hides at position five, that question contributes 0.2.

def reciprocal_rank(retrieved: list[str], relevant: set[str]) -> float:
    for i, doc_id in enumerate(retrieved, start=1):
        if doc_id in relevant:
            return 1.0 / i
    return 0.0

def mrr(results: list[list[str]], golden_sets: list[set[str]]) -> float:
    rrs = [reciprocal_rank(r, g) for r, g in zip(results, golden_sets)]
    return sum(rrs) / len(rrs)

MRR is the right headline metric when your product surfaces retrieved chunks to users directly — search-style UIs, citations the user might click. Position one matters qualitatively more than position five, and MRR encodes that.

NDCG: When Multiple Documents Matter

Normalized Discounted Cumulative Gain generalizes to multiple relevant documents and graded relevance: documents gain credit the higher they rank, and later positions are discounted logarithmically. With binary relevance it still differs from MRR — it accumulates credit for every relevant hit, not just the first. When your prompt consumes six chunks and two of them contain complementary parts of the answer, NDCG is the metric that reflects that.

import math

def dcg(retrieved: list[str], relevance: dict[str, int], k: int) -> float:
    return sum(
        relevance.get(doc_id, 0) / math.log2(i + 1)
        for i, doc_id in enumerate(retrieved[:k], start=1)
    )

def ndcg(retrieved: list[str], relevance: dict[str, int], k: int) -> float:
    best = dcg(  # perfect ranking = ideal DCG
        [d for d, _ in sorted(relevance.items(), key=lambda x: -x[1])],
        relevance, k,
    )
    if best == 0:
        return 0.0
    return dcg(retrieved, relevance, k) / best

The ideal DCG computation here uses the same helper on a perfectly sorted list — a common shortcut, and a correct one as long as the relevance dictionary is built from the same graded labels. If you only have binary labels, the code simplifies considerably, and NDCG degenerates toward a position-weighted multi-hit measure — still more informative than recall alone.

Generation Metrics: Groundedness and Answer Relevance

Retrieval metrics need labels. The generation side has a cheaper instrument: LLM-as-judge scoring against the retrieved context. The two questions that matter are:

  • Faithfulness / groundedness — is every claim in the answer supported by the retrieved context? This catches hallucination that retrieval cannot: the right document came back, and the model embellished anyway.
  • Answer relevance — does the answer actually address the question, rather than being merely on-topic? A perfectly grounded non-answer passes faithfulness and still fails the user.

Both are checkable with a strong model given a strict rubric and structured output. The Ragas library packages exactly these evaluations (its faithfulness metric decomposes the answer into claims and verifies each against the context), and building it yourself is a reasonable learning exercise — but start with the library to avoid rubric mistakes that silently bias your scores.

import json

FAITHFULNESS_PROMPT = """You are verifying a support answer.

Question: {question}
Retrieved context:
{context}
Answer: {answer}

List every factual claim in the answer as a JSON array of strings.
For each claim, decide if it is fully supported by the retrieved context.
Respond with JSON: {{"claims": [{{"text": "...", "supported": true|false}}]}}"""

def faithfulness(answer: str, context: str, question: str, judge) -> float:
    raw = judge(FAITHFULNESS_PROMPT.format(
        question=question, context=context, answer=answer))
    data = json.loads(raw)
    claims = data.get("claims", [])
    if not claims:
        return 1.0
    supported = sum(1 for c in claims if c.get("supported"))
    return supported / len(claims)

Judge-based scores are noisy at the individual level and stable in aggregate. Run every golden question, compare distributions before and after a change, and never draw conclusions from a handful of samples. Also anchor the judge periodically against human labels on a small subset — judge models drift, and your sense of “supported” should not.

Putting the Harness Together

The whole evaluation loop fits in one script that runs against any RAG endpoint exposing a retrieve-and-answer function:

def mean(xs: list[float]) -> float:
    return sum(xs) / len(xs) if xs else 0.0

def evaluate(rag_pipeline, golden_set: list[dict], k: int = 5) -> dict:
    recalls, rrs, faith = [], [], []
    for item in golden_set:
        docs, answer = rag_pipeline(item["question"])
        doc_ids = [d["id"] for d in docs]
        recalls.append(recall_at_k(doc_ids, item["relevant_ids"], k))
        rrs.append(reciprocal_rank(doc_ids, item["relevant_ids"]))
        if "expected_answer" not in item:
            continue
        context = "\n\n".join(d["text"] for d in docs)
        faith.append(faithfulness(
            answer, context, item["question"], judge=call_judge_model))

    report = {
        "recall@%d" % k: mean(recalls),
        "mrr": mean(rrs),
        "questions": len(golden_set),
    }
    if faith:
        report["faithfulness"] = mean(faith)
    return report

Wire this into CI on the changes that silently move quality: embedding-model upgrades, chunk-size changes, reranker swaps, prompt edits. Any of these can shift recall@5 by ten points while every individual answer in manual testing still looks fine. Failing the build on a retrieval-metric drop is the difference between a RAG system and a demo.

Diagnosing With the Numbers

The metrics form a decision tree. Low recall@k means the retrieval layer is the bottleneck — fix chunking, embeddings, or add hybrid keyword search before touching the prompt. Good recall but low MRR or NDCG means ranking is the problem — a cross-encoder reranker over the top fifty candidates is the standard fix. Good retrieval but low faithfulness means the generation side leaks — tighten the prompt’s grounding instructions, reduce the number of chunks so the model has less to blend, or switch to a model with stronger instruction-following. Each symptom has a different remedy, and without the metrics you are guessing which one you have.

Wrapping Up

A minimal but honest RAG evaluation needs three things: a hundred labeled questions, recall@k and MRR computed over them, and a faithfulness check on the answers. That is a day of work, it runs in CI, and it converts “the chatbot seems worse today” into a number with a direction. The Ragas docs cover the judge-based metrics in depth, and the retrieval metrics above come from the classic information-retrieval evaluation measures — half a century old and still the right tools for the job.

Leave a Reply

Your email address will not be published. Required fields are marked *