Start here. This is the direct spoken answer to practice first.
Overview
Reserved capacity is useful only when traffic and reliability needs justify paying for it.
Pay-as-you-go inference uses shared capacity and fits prototypes, variable traffic, and workloads that can tolerate throttling or retry. Provisioned throughput reserves capacity for more predictable performance and is better for sustained production demand with clear latency or availability needs. I compare cost at the real token mix and peak traffic, not only at average requests per minute.