Pick a focused question that fits your time, stack, and interview goal.
How much time do you have?
Show one-drill sessions you can finish now.
40 results across 1 active filter
Page 2 of 2
Builds system-specific adversarial testing around realistic assets, identities, channels, tools, and measurable impact.
Treats model changes as controlled releases using versioned configuration, fixed evaluations, shadow or canary traffic, monitoring, and rollback.
Combines deterministic invariants, severity-aware quality thresholds, slice protection, baseline comparison, and guarded rollout.
Selects a production experiment based on user exposure, side effects, measurement needs, cost, and rollback risk.
Instruments an AI workflow with correlated stage spans, low-cardinality versions, usage and outcome signals, and explicit content-capture controls.
Combines quality, reliability, cost, security, privacy, and recovery evidence into a realistic launch and incident game day.
Diagnoses schema-valid quality regressions across inputs, context, prompts, examples, models, schemas, tools, and downstream interpretation.
Defines hallucination as unsupported or incorrect model output and connects it to evidence, task design, evaluation, and user-visible uncertainty.
Defines AI evaluation as repeatable measurement of application behavior while preserving deterministic tests for code and contracts.
Uses a more precise second-stage ranker on a bounded candidate set while accounting for latency, cost, and authorization.
Uses task bounds and evaluation evidence to choose smaller models for lower latency, cost, capacity, privacy, or local execution.
Matches exact checks, model judgment, and human expertise to the behavior being measured instead of treating graders as interchangeable.
Uses pairwise judgments for relative change and absolute criteria for release obligations while controlling order and verbosity bias.
Balances task quality, safety, reliability, latency, tokens, cost, refusals, repairs, and fallbacks by route and slice.
Separates structural validity from factual accuracy, evidence, authorization, and domain invariants.
Connects probabilistic decoding and changing service conditions to output variation, reproducibility limits, and appropriate testing.