Skill area
Measuring task quality, failures, regressions, and release readiness with representative evidence.
Practice items tagged with AI Evaluation.
Diagnoses an AI regression by decomposing request stages, configuration changes, token growth, retries, routing, tools, retrieval, caching, and provider behavior.
Explains how decoding settings reshape token selection and why they must be tuned against task-specific evaluation rather than folklore.
Balances AI quality, tail latency, and cost through explicit product thresholds, model choice, context control, caching, routing, and measurement.
Chooses a model from task-specific evaluation, operating constraints, safety, latency, cost, context, modality, and provider requirements.
Designs a bounded and reversible first AI feature with measurable value, controlled data, validation, evaluation, observability, fallback, and human oversight.
Treats model changes as controlled releases using versioned configuration, fixed evaluations, shadow or canary traffic, monitoring, and rollback.
Preparation paths where this taxonomy appears in the track scope.