Pick a focused question that fits your time, stack, and interview goal.
How much time do you have?
Show one-drill sessions you can finish now.
148 results across 1 active filter
Page 4 of 7
Judges claims against allowed evidence, separates correctness from support, and tests whether the system answers or abstains appropriately.
Builds a query-and-relevance dataset and separates retrieval recall, ranking quality, context quality, and generated-answer quality.
Measures schema validity, tool selection, argument correctness, policy, execution outcome, repair, and final user result.
Treats the model-facing schema and downstream consumer contract as a versioned API with compatibility and rollout concerns.
Connects answer claims to retrieved evidence while preventing decorative citations and preserving source-level traceability.
Controls admission, concurrency, queues, tenant fairness, token budgets, and overload behavior before provider throttling cascades.
Combines application authorization, scoped platform resources, tenant budgets, admission control, usage attribution, and noisy-neighbor protection.
Coordinates source changes, versions, tombstones, reconciliation, and query-time freshness without treating the index as authoritative.
Keeps credentials in trusted executors and replaces broad reusable secrets with scoped references and short-lived authority.
Uses stable operation identities, persisted states, deduplication, side-effect keys, and replay rules across requests and workers.
Separates durable application state from the model's bounded request context while preserving ownership, concurrency, and deletion.
Rebuilds and compares incompatible vector spaces using versioned indexes, shadow queries, atomic cutover, and rollback.
Applies purpose limitation, selective retrieval, redaction, pseudonymization, and field-level policy before model calls.
Constrains source-to-sink combinations with least privilege, egress policy, argument checks, approval, and information-flow awareness.
Carries trusted tenant scope through retrieval, caches, memory, tools, traces, evaluations, and asynchronous execution.
Uses infrastructure and configuration as code, immutable behavior versions, environment-specific resources, evaluation gates, and drift detection.
Checkpoints authoritative state, reconciles ambiguous tool outcomes, preserves approvals, and resumes without repeating completed work.
Builds system-specific adversarial testing around realistic assets, identities, channels, tools, and measurable impact.
Uses constrained output where possible, deterministic parsing and validation, bounded targeted repair, and safe fallback.
Contains compromised AI capabilities, preserves evidence, scopes derived data and actions, remediates boundaries, and verifies safe recovery.
Uses explicit capability and risk policy to route by task while preserving evaluation, version attribution, budgets, and fallback semantics.
Tracks trust, versions, integrity, provenance, review, and rollback across models, adapters, prompts, datasets, tools, and parsers.
Establishes server trust, audience-bound authorization, least privilege, capability filtering, approval, and safe handling of untrusted results.
Uses workload identity, separate control and data planes, least privilege, scoped model access, and auditable administration.