
Arjun Sethi
I start by arguing that retrieval-augmented generation has three independent stages, and most disappointing results come from improving the wrong one. The stages are retrieval, reranking and generation. Retrieval finds candidate passages. Reranking reorders those candidates with a model that reads the query against each passage. Generation answers from what survived. If retrieval never surfaces the right passage, the model cannot recover, and no amount of prompt work fixes that. Chunking is the decision with the most downstream effect. Fixed-size chunks cut sentences in half. Semantic chunks follow content boundaries and produce variable sizes. Recursive splitting tries to respect a size limit while preserving structure. Every one needs overlap, because a fact split across a boundary is invisible to both halves, and overlap large enough to cover the longest relevant passage costs context on every query. Reranking earns its cost at the top of the results. Returning five passages and reranking them is usually better than returning twenty and hoping. Evaluation is where I spend the most page space. Recall at k tells you whether the answer was in the results at all, and it is the number that separates a retrieval problem from a generation problem. Ranking metrics tell you whether it appeared near the top. Both are cheap to compute against a small hand-labelled set, and without one you are tuning blind. The context budget matters too. More retrieved text does not mean a better answer, and a model given twelve mediocre passages will use twelve mediocre passages.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →