Start here. This is the direct spoken answer to practice first.
Overview
An aggregate can improve while a small but important class of requests becomes much worse.
A slice is a meaningful subset of evaluation cases, such as language, task type, customer tier, risk level, input length, or tool path. I report performance by slices that reflect supported behavior and known failure modes, not only one overall score. This reveals regressions that common easy requests would otherwise dominate.