Start here. This is the direct spoken answer to practice first.
Overview
An eval stops measuring generalization when the system is repeatedly optimized against the exact visible cases.
I separate cases used during prompt development from a held-out set used for comparison, and I limit who or what can inspect the held-out answers. I check for duplicates and near-duplicates across training, tuning, and evaluation data where those sources are known. A change must improve representative slices, not merely the cases that motivated it.