Start here. This is the direct spoken answer to practice first.
Overview
AI cost and capacity can be attacked through more dimensions than request count alone.
I limit request size, files, context, output tokens, tool calls, retries, agent turns, fan-out, concurrency, and elapsed time per user, tenant, and workload. Admission control reserves capacity for important traffic and rejects or queues work before expensive execution. Every run has a budget the model cannot increase.