Start here. This is the direct spoken answer to practice first.
Why this question matters
Many incidents are not instant failures; they are saturation building up until the system cannot recover by itself.
I would watch leading signals: CPU, memory, thread pool starvation, connection pool usage, SQL waits, Redis latency, queue age, consumer lag, request concurrency, timeout rate, retry rate, and dependency saturation. User-facing latency and errors are lagging symptoms. Leading signals tell me the system is approaching a cliff before users fall off it.