Start here. This is the direct spoken answer to practice first.
Overview
A model is appropriately sized when it meets the task requirement reliably, not when it has the highest general benchmark score.
I use a smaller model when the task is narrow and an evaluation shows it meets the required quality. Classification, routing, extraction into a constrained schema, moderation prechecks, or rewriting short text may not need the strongest general model. Smaller models can reduce latency and cost, increase throughput, and sometimes run in a controlled or local environment. I do not choose one from size alone; I compare it on the feature's actual inputs.