Pick a focused question that fits your time, stack, and interview goal.
How much time do you have?
Show one-drill sessions you can finish now.
63 results across 1 active filter
Page 1 of 3
Uses traces and saturation evidence to locate retry amplification, queue growth, exhausted workers, and failing fallback paths.
Uses traces to separate unclear goals, bad tool interfaces, missing state, contradictory results, poor termination, and retry amplification.
Diagnoses an AI regression by decomposing request stages, configuration changes, token growth, retries, routing, tools, retrieval, caching, and provider behavior.
Reproduces failing cases, compares versioned traces, isolates the first changed component, mitigates impact, and adds regression coverage.
Traces a missed answer through source ingestion, parsing, chunking, filtering, query construction, retrieval, fusion, and ranking.
Builds a reproducible evaluation harness that records immutable cases, versioned outputs, deterministic checks, calibrated graders, slices, and release decisions.
Builds a schema-constrained extraction boundary that distinguishes refusal, malformed output, semantic invalidity, and safe bounded repair.
Balances burst flexibility, predictable capacity, latency variance, commitment cost, and realistic traffic shape.
Handles a poisoned retrieval corpus by freezing ingestion, tracing provenance, switching immutable index versions, rebuilding clean data, and proving recovery.
Connects product success, layered evals, calibrated graders, release gates, guarded rollout, traces, monitoring, feedback, and incident learning.
Combines explicit trust boundaries, provider access, state, execution modes, resilience, budgets, versioning, and operations into a restrained design.
Uses workflow traces to distinguish planning failure, stale observations, retry ambiguity, and missing terminal conditions in a looping agent.
Uses versioned traces and evaluation slices to locate whether a RAG regression came from ingestion, retrieval, context assembly, or generation.
Breaks results down by meaningful user, task, risk, language, input, and system dimensions instead of trusting one average.
Separates development and held-out cases, controls access, detects duplication, and validates improvements on fresh traffic.
Balances AI quality, tail latency, and cost through explicit product thresholds, model choice, context control, caching, routing, and measurement.
Handles token-budget overflow and incomplete generation through explicit prioritization, rejection, retrieval, summarization, continuation, and validation.
Separates provider refusals, safety blocks, token truncation, transport failure, and valid domain abstention.
Separates model and provider outcomes, applies bounded recovery, validates outputs, and exposes useful status and telemetry.
Uses deadline budgets, transient-only bounded retries, jitter, idempotency, and propagated cancellation without retry storms.
Records attributable versions, decisions, evidence references, policy results, and side-effect receipts while minimizing content-rich logs.
Defines success, failure, step, time, token, cost, retry, and repetition limits with useful escalation behavior.
Combines production-shaped, expert-authored, edge, and adversarial cases with versioned labels and privacy controls.
Chooses among shared, provisioned, regional, data-zone, global, batch, or managed-compute patterns from workload and governance constraints.