Pick a focused question that fits your time, stack, and interview goal.
How much time do you have?
Show one-drill sessions you can finish now.
40 results across 1 active filter
Page 1 of 2
Diagnoses an AI regression by decomposing request stages, configuration changes, token growth, retries, routing, tools, retrieval, caching, and provider behavior.
Reproduces failing cases, compares versioned traces, isolates the first changed component, mitigates impact, and adds regression coverage.
Builds a reproducible evaluation harness that records immutable cases, versioned outputs, deterministic checks, calibrated graders, slices, and release decisions.
Connects product success, layered evals, calibrated graders, release gates, guarded rollout, traces, monitoring, feedback, and incident learning.
Combines governed ingestion, authorized hybrid retrieval, reranking, grounded generation, citations, evaluation, observability, and fallback.
Uses workflow traces to distinguish planning failure, stale observations, retry ambiguity, and missing terminal conditions in a looping agent.
Uses versioned traces and evaluation slices to locate whether a RAG regression came from ingestion, retrieval, context assembly, or generation.
Breaks results down by meaningful user, task, risk, language, input, and system dimensions instead of trusting one average.
Explains how decoding settings reshape token selection and why they must be tuned against task-specific evaluation rather than folklore.
Separates development and held-out cases, controls access, detects duplication, and validates improvements on fresh traffic.
Balances AI quality, tail latency, and cost through explicit product thresholds, model choice, context control, caching, routing, and measurement.
Combines production-shaped, expert-authored, edge, and adversarial cases with versioned labels and privacy controls.
Chooses a model from task-specific evaluation, operating constraints, safety, latency, cost, context, modality, and provider requirements.
Tunes evidence quantity and no-answer behavior from measured relevance rather than treating similarity scores as universal confidence.
Combines low-friction ratings, structured reasons, behavioral outcomes, review queues, and privacy-aware eval case creation.
Turns a product outcome into observable quality, safety, and operational criteria before selecting convenient scores.
Builds a focused scoring guide, representative labeled examples, bias checks, and ongoing agreement monitoring for model-based evaluation.
Designs a bounded and reversible first AI feature with measurable value, controlled data, validation, evaluation, observability, fallback, and human oversight.
Separates traffic, data, dependency, model, prompt, and grader change while maintaining stable anchors and fresh evaluation.
Evaluates verified outcomes, trajectory quality, safety, efficiency, recovery, and important slices in realistic environments.
Judges claims against allowed evidence, separates correctness from support, and tests whether the system answers or abstains appropriately.
Builds a query-and-relevance dataset and separates retrieval recall, ranking quality, context quality, and generated-answer quality.
Measures schema validity, tool selection, argument correctness, policy, execution outcome, repair, and final user result.
Connects answer claims to retrieved evidence while preventing decorative citations and preserving source-level traceability.