Start here. This is the direct spoken answer to practice first.
Overview
The most reliable eval suites combine evaluator types because each one sees a different class of failure.
I use deterministic checks for facts the application can verify exactly, such as schema validity, required citations, executable code, or permitted tool arguments. I use a model grader for scalable judgments with a clear scoring guide, such as relevance or whether a response addresses the request. Human experts remain necessary for ambiguous, high-impact, or newly discovered cases and to verify that automated graders still agree with the intended standard.