← 🤖 AI Infrastructure
Core
Why GPUs break the assumptions schedulers are built on
GPU scheduling must account for device memory, sharing boundaries and interconnect topology. Device plugins, MIG, DRA and gang scheduling address different constraints; none makes accelerators behave like interchangeable CPU millicores.
a free account opens the core tier · no card
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS
AI & GPU InfrastructureDevice plugin versus Dynamic Resource Allocation for GPUs: why did Kubernetes need DRA?→AI & GPU InfrastructureYour cluster has three GPU generations because that is what you could buy. How do you schedule against it?→AI & GPU InfrastructureFour teams want GPUs and you have twelve A100s. Walk me through MIG, time-slicing and MPS, and how you would decide.→Containers & KubernetesA pod requesting one GPU stays Pending on a cluster with idle GPU nodes. Debug it.→Incident Response & Production DebuggingA pod has been Pending for ten minutes. Walk me through your diagnosis.→AI & GPU InfrastructureA multi-node training job sits Pending forever while the cluster shows free GPUs. What is happening?→