Retrieval-Augmented Generation

Vector Search vs Agentic Retrieval: When grep Beats a Retriever

Vector retrieval compresses a corpus into one similarity score per passage. Five papers measure the cost and how grep agents, small readers, citation grading, and per-chunk configs each recover a layer.

Vector Search vs Agentic Retrieval: When grep Beats a Retriever

The question every RAG system answers badly

Since the 2020 RAG paper, the default answer to “how does a model get the right evidence?” has been the same: embed the corpus, embed the query, return the top-K passages. That recipe won because it is simple, fast, and works well enough for one-shot open-domain QA. But every design since has attacked the same weakness from a different angle, and the papers collected on this site read as a coordinated critique.

The critique starts from a simple observation: a similarity score is a compression. When a retriever hands an agent a ranked list of whole passages, it has already decided what matters, at a resolution of whole documents, before any reasoning happened. Exact lexical constraints vanish. Conjunctions of weak clues cannot be expressed. Evidence filtered out at the top-K step can never be recovered, no matter how strong the downstream model is. Five papers on this site each fix a different layer of that pipeline: the interface, the reader, the citation, the configuration, and the measurement. None of them says vector search is dead. All of them say the vector index is the wrong place to stop.

What vector search compresses away

The clearest statement of the resolution problem is Direct Corpus Interaction, or DCI. Its authors ran a controlled swap: same Claude Sonnet 4.6 backbone, same tasks, but instead of calling a Qwen3-Embedding-8B retrieval tool, the agent gets a terminal and reads the raw corpus with grep, find, and file reads. Accuracy on BrowseComp-Plus went from 69.0% to 80.0%, an 11-point gain, while evaluation cost fell 29.4% from $1,440 to $1,016. The agent did not win because grep surfaced more gold documents. Ablations show DCI often wins even when the retrieval agent already surfaced all the gold evidence. It wins because it converts coarse evidence into fine-grained verification: pipe two greps together and you have a conjunction; grep for a keyword and read the surrounding lines and you have a local-context check. A top-K list can do neither.

The 2020 RAG paper itself already contains the seed of this critique. Its retriever returns a fixed top-K (around 5 to 10 passages) over 21M Wikipedia chunks, and if DPR misses the relevant passage, the BART-large generator has nothing to ground on. The paper’s own framing treats retrieval as a marginalized latent variable trained end to end, which modern production RAG quietly dropped. What survived is the interface, top-K in, prompt out, and the interface is exactly the bottleneck DCI attacks.

Key numbers

ApproachPaperHeadline numbersBest atBreaks when
Dense retriever plus seq2seq generatorRAG (2020)Top-K of 5 to 10 over 21M Wikipedia chunks; BART-large generator of ~400M params; state of the art on NQ, TriviaQA, and WebQuestions at publicationOne-shot open-domain QA; knowledge updated by swapping the index, no retrainingA retriever miss is unrecoverable; fixed top-K caps recall
Agent greps the raw corpusDCIBrowseComp-Plus 69.0% to 80.0% (+11.0) at 29.4% lower cost ($1,440 to $1,016) under the same Sonnet 4.6 backbone; multi-hop QA 83.0 average (+30.7); NDCG@10 68.5 average (+21.5)Conjunctions of weak clues, exact lexical checks, local verification400K documents: 37.5% accuracy and 122.4 tool calls per question
Tiny context-only readerOCC-RAG0.6B and 1.7B params trained on 3M+ synthetic QA examples; matches or beats general models 2 to 6 times its size on five grounded-QA benchmarksCheap faithful answering with quote-level citations and trained refusalWithout retrieved context it has almost nothing to reason over
Citation-graded document QACiteVQA1,897 questions over 711 PDFs averaging 40.6 pages; Strict Attributed Accuracy 76.0 for Gemini-3.1-Pro vs 22.5 for the best open modelCatching right-answer-wrong-citation failures across 20 MLLMsA benchmark, not a fix; ships no model
Per-chunk retrieval configCARVERecall@5 0.603 vs 0.510 for the best of 8 baselines; generation pass rate 0.357 vs 0.315 with Qwen3-VL-8B; beats trained routers at 0.329 and 0.310Long-video RAG where the right retrieval format is a property of the chunkFour retrievers plus a 2B reranker per query; absolute scores stay low

Five fixes at five layers

The table hides the fact that these papers are not competitors. They are answers to different failure modes, and stacking them is the actual lesson.

The interface. DCI removes the retriever and lets a capable agent search like a researcher. The operating envelope is explicit: at 100K documents the agent makes 38.5 tool calls per question, at 200K it makes 86.9, and accuracy drops 13.6 points. At 400K documents accuracy collapses to 37.5% with 122.4 calls per question and 20 examples running out of tool budget entirely. DCI scales in search depth, not search breadth. The authors themselves agree that dense and sparse retrieval remain the right tool for large static corpora, and DCI is for local, evolving, agent-controlled workspaces.

The reader. OCC-RAG accepts the retriever as a given and asks a different question: what should the reading model be optimized for? Its answer is that parameters spent memorizing world facts are mostly wasted in a grounded system, and sometimes harmful, because they tempt the model to answer from memory instead of from the passage. So it trains 0.6B and 1.7B models on over 3 million synthetic examples engineered for two behaviors: staying faithful to the provided context, and refusing to answer when the context does not support one. The payoff is that a 1.7B model matches or beats general-purpose models two to six times its size on HotpotQA, MuSiQue, TAT-QA, ConFiQA, and MuSiQue-Un, while quoting its evidence verbatim inside the reasoning trace.

The citation. CiteVQA targets the failure that makes RAG dangerous in production: attribution hallucination. Its premise is that in contracts, filings, and medical records, a correct answer with a wrong citation is worse than no answer, because the citation is what a reviewer trusts. So the benchmark grades answers and bounding-box evidence jointly. The result across 20 MLLMs: Gemini-3.1-Pro-Preview tops out at 76.0 Strict Attributed Accuracy, the best open model (Qwen3-VL-235B-A22B) manages 22.5, and the roughly 53-point gap is the measured size of the attribution problem. This is the paper to cite when someone claims their RAG pipeline is “fully grounded.”

The configuration. CARVE shows that even the choice of what to retrieve is usually made at the wrong granularity. In long-video RAG, the common practice fixes one retrieval config (visual embeddings or text summaries, keyframes or full clips) per query. CARVE runs four retrievers in parallel and assigns each chunk the config that retrieved it best. The ablation is clean: the best single config reaches 0.507 Recall@5, the best two-config combination 0.567, and per-chunk selection over all four reaches 0.603 against a 0.510 best baseline. Reranking each chunk only under its own config is load-bearing: random config assignment or joint concatenation both drop to 0.513. And a rule-based per-chunk scheme with no training at all beats a trained non-LLM router (0.329) and a LoRA-tuned LLM router (0.310) on generation pass rate.

The measurement. None of these numbers mean anything if the benchmark leaks. V-RAGBench, the benchmark CARVE ships, exists because more than half of widely used video QA samples are answerable without the video at all, from language priors, world knowledge, or static cues. Its builders filter candidates through five sequential checks, including answerability verification with GPT-5.2-chat and an evidence-uniqueness test, down to 2,100 evidence-labeled triplets over 216 videos. Because evidence is labeled, retrieval and generation are scored separately, which is what makes the 0.603 Recall@5 number interpretable. The same instinct, grade the retrieval itself, not just the final answer, is what CiteVQA applies to citations and what DCI’s controlled swap applies to interfaces.

When to use which

  • Large, static corpus, simple lookup questions. Keep vector retrieval. It is cheap, predictable, and DCI’s own scaling curve shows agentic search paying 122.4 tool calls per question at 400K documents. Add OCC-RAG-style small faithful readers if per-query cost matters and your queries are answerable from retrieved context alone.
  • Local, evolving corpus; queries that combine weak clues. This is DCI’s home turf. Research workspaces, codebases, and document trees where exact strings, conjunctions, and local context checks carry the load. Budget for tool calls, and cap the corpus size or hybridize with a retriever once breadth matters.
  • High-stakes document QA. Add citation grading in the CiteVQA style: require element-level evidence with every answer and score them jointly. Treat a 22.5 SAA from a strong open model as the honest current baseline, not an anomaly.
  • Multimodal or long-video RAG. Do not fix one retrieval config per query. CARVE’s numbers say the config is a property of the chunk; per-chunk selection costs you four retrievers and a reranker but buys 0.093 absolute Recall@5 over the best single config.
  • Any of the above. First check the benchmark for leakage. If more than half the questions are answerable without the corpus, every downstream number is decoration.

Limits and open questions

The honest limit is that these papers do not share a benchmark, so the table above compares numbers that were never run head to head. DCI’s 80.0% is on BrowseComp-Plus with a frontier agent; CARVE’s 0.603 is on V-RAGBench with egocentric video; CiteVQA’s 76.0 ceiling is document QA with bounding boxes. The “same harness?” column that a GRPO vs PPO comparison can fill does not exist here, and anyone claiming a single winner across all three is overselling.

Each approach also has a known soft spot. DCI leans on frontier models (Sonnet 4.6, and a GPT-5.4 nano minimal harness) and on curated BRIGHT and BEIR sets, so part of the gain is the agent’s strength, not just the interface. OCC-RAG’s “matches models 2 to 6 times its size” is a range claim, not an audited per-benchmark margin, and its 3M training examples are synthetic, which risks teaching the style of grounded answering rather than surviving real retrieval noise. CiteVQA diagnoses attribution hallucination but ships no fix, and two of its four metrics rely on an LLM judge. CARVE pays four retrievers plus a 2B reranker per query and reports no cost budget, and its 0.603 Recall@5 still means the right evidence misses the top 5 about 40% of the time.

The open question that ties all five together: who trains the interface? The 2020 RAG paper trained retriever and generator jointly and the industry dropped that training signal for engineering convenience. DCI and CARVE recover capability by spending inference-time compute instead. Whether jointly trained retrieval comes back, or whether agentic interfaces make the retriever fully replaceable, is genuinely unsettled.

FAQ

On the tested setups, yes, with caveats. Under the same Claude Sonnet 4.6 backbone, replacing a Qwen3-Embedding-8B retrieval tool with DCI’s terminal access raised BrowseComp-Plus accuracy from 69.0% to 80.0% while cutting cost 29.4%. On multi-hop QA the same setup averaged 83.0, 30.7 points over the strongest retrieval-agent baseline. But the comparison is interface-versus-interface at frontier model quality on curated sets, not a universal verdict, and it reverses as corpus breadth grows.

When does DCI retrieval stop scaling?

Fast, and measurably. Tool calls per question rise from 38.5 at 100K documents to 86.9 at 200K, accuracy falls 13.6 points, and at 400K documents accuracy is 37.5% with 122.4 calls per question and 20 examples out of budget. DCI scales well in search depth and poorly in search breadth, which is why its authors point large static corpora back to dense and sparse retrieval.

Why do RAG answers come out right with citations that are wrong?

Because answer-only benchmarks never graded the citation. CiteVQA scores the answer and the bounding-box evidence jointly at IoU at least 0.5, and across 20 MLLMs the pattern is systematic: models answer correctly while pointing at the wrong region. That is attribution hallucination, and it is why Strict Attributed Accuracy tops out at 76.0 for Gemini-3.1-Pro-Preview while the best open model, Qwen3-VL-235B-A22B, sits at 22.5.

What does CARVE’s per-chunk retrieval config add over query-level routing?

About 0.093 absolute Recall@5 over the best baseline, and it beats learned routing while doing it. The best single config reaches 0.507, the best two-config combination 0.567, and per-chunk selection over all four configs reaches 0.603 on V-RAGBench. A trained non-LLM router scores 0.329 and a LoRA-tuned LLM router 0.310 on generation pass rate, both below CARVE’s untrained per-chunk rule at 0.357.

Is a 0.6B faithful reader enough for production RAG?

Enough for the reading half, if your retrieval is good. OCC-RAG’s 0.6B and 1.7B models match or beat general models two to six times their size on five grounded-QA benchmarks, and they abstain when the context does not support an answer. But they are deliberately built to be useless without retrieved context, so the retriever upstream is still your problem, and the faithfulness claim is a range measured against synthetic training data.