Start here. This is the direct spoken answer to practice first.
Why this question matters
Alerting is easy to overdo and hard to make useful. This drill matters because the production behavior is shaped by platform configuration, code boundaries, and the operational path the team will use during a real incident.
I would alert on sustained symptoms that reflect user impact: availability failures, high error rate, high latency, unhealthy instances, critical dependency failures, and queue age for background work. I would avoid paging on every isolated exception because that creates noise and teaches the team to ignore alerts. Each alert needs a clear owner and a first action.