Start here. This is the direct spoken answer to practice first.
Overview
An evaluation result is only as relevant as the inputs, labels, and slices behind it.
I begin with real task categories and production-shaped inputs, then add expert-authored cases for important behavior that is rare or not yet observed. The dataset includes ordinary requests, ambiguous inputs, edge cases, unsupported requests, and adversarial variations. Each case records its source, expected behavior, applicable scoring guide, and important slice labels so results can be interpreted rather than treated as one anonymous average.