Start here. This is the direct spoken answer to practice first.
Overview
More retrieved chunks can add evidence, but they can also add noise, contradiction, latency, and cost.
I tune top-k on a labeled query set by measuring whether the needed evidence appears and how much irrelevant context reaches generation. Similarity scores are ranking signals, not universal probabilities, so a threshold must be calibrated for the chosen model, index, distance metric, and query type. When evidence is missing or weak, the application asks for clarification, falls back to search, or says it cannot answer.