Start here. This is the direct spoken answer to practice first.
Overview
Safety behavior must match the product's users, domain, actions, and harm model rather than a generic blocked-word list.
I begin with a product policy that defines disallowed, restricted, and allowed uses for the actual audience and domain. Inputs, retrieved content, generated outputs, and consequential actions can each need different checks. The system returns a safe explanation or escalation path rather than leaking hidden policy or failing unpredictably.