Start here. This is the direct spoken answer to practice first.
Overview
A useful RAG record needs enough identity and policy metadata to be governed after ingestion.
I store a stable chunk ID, document ID, source version, source URI or locator, section path, offsets, content hash, timestamps, language, and the parser, chunker, and embedding versions. Filterable metadata includes tenant, document type, status, and access principals where required. The searchable text and vector are never the only representation of a source.