Start here. This is the direct spoken answer to practice first.
Overview
The investigation should locate which stage and request slice changed before attempting prompt or model optimizations.
I confirm the time window and affected feature, tenant, model, and request class, then split latency and cost by stage: context construction, retrieval, model calls, tool calls, validation, retries, and queueing. I compare the deployed configuration and input distributions with the previous baseline. Common causes include longer history or retrieved text, a model or routing change, more generated tokens, retry storms, cache misses, slower tools, or provider throttling.