Published 6 October 2026
Most RAG systems work fine in the demo and then quietly disappoint in production. The pattern is familiar: a proof of concept answers a handful of test questions well, gets approved, ships, and then three weeks later someone notices the system is confidently citing the wrong policy clause or missing a document that’s clearly relevant. The model isn’t usually the problem. The pipeline feeding it context is.
Retrieval-augmented generation looks simple on a whiteboard — embed documents, store vectors, retrieve the closest matches, hand them to the model. In practice, almost every step in that chain has a failure mode that doesn’t show up until the system meets real documents and real queries. Here’s where that actually goes wrong, and what fixing it looks like.
Chunking Strategy Is Usually the First Mistake
Fixed-size chunking — splitting documents every 500 or 1000 characters regardless of structure — is the default in most tutorials and the first thing that breaks on real content. A contract clause, a table row, or a procedure step gets sliced in half, and the retriever ends up with two chunks that are each individually meaningless. The model then either hallucinates a connection between unrelated fragments or answers from incomplete context without flagging that it’s incomplete.
The fix is chunking that respects document structure — splitting on headings, paragraphs, table boundaries, or semantic units rather than raw character counts — and it has to be done per document type rather than as one global rule. A technical manual, a legal contract, and a support ticket thread all have different natural units, and a single chunking strategy tuned for one will quietly degrade the others.
Embedding Model Mismatch
Teams frequently pick an embedding model based on a generic benchmark leaderboard and never revisit that choice once it’s in production. Embedding models vary significantly in how well they capture domain-specific meaning — a model tuned on general web text may not distinguish between closely related technical terms in a specialized domain, which means queries that should retrieve distinct documents end up retrieving nearly identical similarity scores for both.
This matters more in dense technical or regulatory domains than in general knowledge-base use cases, and it’s worth testing embedding models against your actual document set and query patterns before committing, not just against a public benchmark. A model that’s mediocre on general retrieval benchmarks can outperform a leaderboard leader on your specific corpus, and the only way to know is to test both against real queries from your domain.
Retrieval Depth: Too Few or Too Many Chunks
Retrieving too few chunks starves the model of context it needs; retrieving too many buries the relevant passage in noise and increases the odds the model latches onto an irrelevant but superficially similar chunk instead. There’s no universal right number — it depends on chunk size, document density, and how often a correct answer genuinely requires synthesizing multiple sources versus one clear passage — but most systems default to a fixed top-k without ever testing whether that number actually serves the query distribution they see in practice.
A more reliable approach is dynamic retrieval depth tied to a relevance-score threshold rather than a fixed count, combined with periodic review of queries where the system retrieved confidently but answered wrong — that failure pattern usually points directly at a retrieval depth or chunking problem rather than a generation problem.
No Evaluation Set, So No Way to Know It’s Broken
This is the pitfall underneath most of the others: teams ship RAG systems without a held-out set of representative queries and known-correct answers to test against, which means regressions are invisible until a user complains. Without an evaluation set, every change to chunking strategy, embedding model, or prompt template is a guess rather than a measured improvement, and it’s impossible to tell whether a “fix” actually helped or just moved the failure pattern somewhere else.
Building even a modest evaluation set — fifty to a few hundred representative query-answer pairs drawn from real usage or domain expert review — before scaling a RAG system past the pilot stage pays for itself almost immediately, because it turns every subsequent tuning decision from a guess into a measurement.
Stale or Duplicated Source Content
Knowledge bases drift. Policies get updated, old versions don’t get removed from the index, and the retriever has no way to know which version is current — it just returns whichever chunk scores highest on similarity, which is sometimes the outdated one. This is less a modeling problem than a content-lifecycle problem, and it’s one that gets worse the longer a RAG system runs without a defined process for re-indexing, deduplicating, and retiring stale source documents.
Systems built for enterprise knowledge retrieval need an explicit content pipeline — not just an initial ingestion job — that handles versioning and removes superseded documents from the index on a defined schedule, otherwise the retrieval quality degrades silently as the underlying knowledge base ages.
Ignoring the “No Good Answer” Case
A RAG system asked a question with no good answer in its knowledge base will still retrieve the closest available chunks and the model will often still generate a confident-sounding response from them, because nothing in the architecture forces it to recognize that the retrieved context doesn’t actually answer the question. This is one of the more damaging failure modes because it looks identical to a correct answer until someone checks the source.
Handling this well requires an explicit confidence or relevance check between retrieval and generation — a step that evaluates whether the retrieved chunks are actually relevant enough to answer the query before passing them to the model, and routes to an “I don’t have enough information” response rather than forcing an answer when they aren’t.
Treating RAG as a One-Time Build Instead of a Maintained System
The pitfalls above share a common root cause: treating RAG implementation as a project with a finish line rather than a system that needs ongoing tuning as the document set, query patterns, and underlying models change. Embedding models improve, document volumes grow, user query patterns shift as adoption increases — a RAG system tuned once at launch and never revisited will degrade over time even if nothing about the implementation was wrong on day one.
The teams that get the most durable value from RAG treat it the way they’d treat any production system with monitoring and iteration cycles — periodic evaluation re-runs, retrieval quality spot-checks tied to real user feedback, and a defined process for re-indexing as source content changes — rather than a one-time integration that’s expected to keep working indefinitely without attention.
Where to Start Fixing an Underperforming System
If a RAG system already in production is underperforming, the highest-leverage first step is almost always building the evaluation set that should have existed from the start, because it’s the only way to diagnose which of the pitfalls above is actually responsible rather than guessing. From there, chunking strategy and retrieval depth tend to produce the largest improvements relative to effort, with embedding model changes and content-lifecycle fixes following once the measurement is in place to confirm they’re helping.
If you’re scoping a new RAG build or trying to diagnose why an existing one underperforms, start a project conversation and we’ll walk through where your specific implementation is likely losing retrieval quality.
Related
