Start here. This is the direct spoken answer to practice first.
Overview
When the schema still passes, the defect is likely in task meaning, evidence, model behavior, or downstream interpretation rather than serialization.
I first define the user-visible regression and compare failing examples with a known-good period. Then I check changes to the model, prompt, examples, schema descriptions, context builder, tools, and downstream rules. I reproduce cases using recorded configuration and evaluate semantic correctness, not only parse success.