Start here. This is the direct spoken answer to practice first.
Overview
End-to-end answer scores alone cannot tell whether retrieval or generation caused a failure.
I create representative queries with relevant document or chunk judgments, including valid no-answer cases. Retrieval metrics such as recall at k, precision at k, mean reciprocal rank, or normalized discounted cumulative gain measure whether useful evidence was found and ordered. Generation is then evaluated on answer correctness, support, completeness, and citation behavior using a fixed retrieved context.