Start here. This is the direct spoken answer to practice first.
Overview
A judge model is another probabilistic component and needs its own evaluation.
I give the judge one focused question, explicit criteria, the relevant input and evidence, and a constrained score such as pass or fail with a reason. I compare its decisions against expert-labeled examples before using it at scale. Calibration checks overall agreement and the kinds of cases where the judge disagrees, because a good average can still hide a systematic bias.