Topic
Representative datasets, task metrics, graders, release comparisons, regressions, and feedback loops.
Practice items tagged with AI Evaluation.
Diagnoses an AI regression by decomposing request stages, configuration changes, token growth, retries, routing, tools, retrieval, caching, and provider behavior.
Explains how decoding settings reshape token selection and why they must be tuned against task-specific evaluation rather than folklore.
Balances AI quality, tail latency, and cost through explicit product thresholds, model choice, context control, caching, routing, and measurement.
Chooses a model from task-specific evaluation, operating constraints, safety, latency, cost, context, modality, and provider requirements.
Designs a bounded and reversible first AI feature with measurable value, controlled data, validation, evaluation, observability, fallback, and human oversight.
Treats model changes as controlled releases using versioned configuration, fixed evaluations, shadow or canary traffic, monitoring, and rollback.