Start here. This is the direct spoken answer to practice first.
Overview
Model selection is a product and operating decision supported by a task-specific evaluation, not a one-dimensional benchmark contest.
I start with the feature's real inputs, expected outputs, failure cost, and latency and budget constraints. Then I build a representative evaluation set and compare candidate models on the qualities the feature needs: correctness, instruction following, structured output, context size, language or modality support, refusal behavior, latency, and cost. I choose the smallest or simplest model that meets the requirement with acceptable margin rather than automatically choosing the most capable model.