WHY I’M THINKING ABOUT THIS

Proofline pushed me to stop treating every wrong answer as a model problem. If the right evidence never reaches the model, a better prompt can still produce a beautifully written wrong answer. The debugging path has to begin earlier.

THE SIGNAL

When a RAG application gives a bad answer, teams often start by changing the prompt or replacing the model.

Sometimes the model never had the right evidence.

Anthropic’s Contextual Retrieval experiments make the point unusually clearly. In its test setup, the relevant-document top-20 retrieval failure rate started at 5.7%. Contextual embeddings reduced it to 3.7%; adding contextual BM25 reduced it to 2.9%; adding reranking reduced it further to 1.9%.

Those are Anthropic’s experimental results—not a universal benchmark.

But they illustrate an important debugging principle:

RAG quality has a pipeline.

THE QUESTION

When the user says “the AI was wrong,” where did the error actually happen?

A simplified failure chain is:

Source → parsing → chunking → indexing → retrieval → filtering → reranking → context → generation → citation → policy

If a policy document is stale, the model can faithfully repeat stale information.

If the correct API error code never appears in the retrieved context, a stronger prompt cannot recover it reliably.

If the right document is retrieved but the answer invents an unsupported claim, that is a generation/faithfulness problem.

These are different product failures and require different fixes.

THE OBVIOUS ANSWER

“Use a better model.”

Sometimes yes.

But Microsoft’s RAG evaluation documentation explicitly separates retrieval/process evaluation from response-level measures such as groundedness and relevance.

That separation is essential.

I would never want a single “RAG accuracy” number.

THE TENSION

A production support system needs multiple eval layers:

Layer What I would measure
Corpus freshness, version correctness, ownership
Retrieval recall@K / NDCG / exact-code retrieval
Reranking most-applicable evidence reaches top context
Generation faithfulness, completeness, unsupported claims
Citation cited source actually supports claim
Safety PII, prompt injection, scope/permission
Product resolution, repeat contact, escalation quality

OWASP also warns that RAG does not eliminate prompt-injection risk. Retrieved content itself can become an attack surface, which means “grounded” is not automatically “safe.”

RAG can fail before the LLM even starts answering supporting data visual
Secondary-research visual. Source and interpretation are stated inside the chart and article.

MY PRODUCT TAKE

For high-stakes support, I would explicitly separate:

Can I find relevant evidence?

from

Can I answer from that evidence?

from

Am I allowed to answer this question?

That third question matters when a user asks for account-specific state that public documentation cannot provide.

A good refusal can be a successful product outcome.

WHAT I WOULD TEST

I would build an eval set with:

Then I would classify failures by pipeline stage before making any model change.

WHAT WOULD CHANGE MY MIND

If the corpus is small enough to fit reliably into context, a complex RAG pipeline may be unnecessary. Anthropic itself notes that, for sufficiently small knowledge bases, passing the corpus directly can be simpler.

The point is not to use RAG.

The point is to use the simplest evidence architecture that produces trustworthy answers.

Secondary research