Start here. This is the direct spoken answer to practice first.
Overview
A metric is useful only after the team agrees what a successful user outcome and an unacceptable failure look like.
I start with the user task, the decision the system may influence, and the failures that matter. For a ticket-triage assistant, success might mean choosing the correct queue, preserving priority, and escalating uncertain cases instead of merely producing plausible prose. I then choose separate measures for task correctness, unsupported behavior, safety, latency, and cost rather than compressing all of them into one score.