Start here. This is the direct spoken answer to practice first.
Overview
A model upgrade is a behavior change, so a successful API call is not enough evidence to promote it.
I version the current model, prompt, tools, retrieval settings, and decoding configuration, then run the candidate against the same representative evaluation set. I compare task quality, important failure slices, invalid outputs, refusals, latency, and cost. If the candidate clears the gate, I use shadow traffic where privacy and cost allow, or a small canary, and monitor both technical and task metrics. The previous configuration remains available for fast rollback.