← 🤖 AI Infrastructure
Advanced
GPU memory allocation: separate live tensors, cached blocks and fragmentation
Diagnose GPU allocation failures using live allocation, reserved memory and allocator evidence. Work through a block-allocation model without assuming every OOM is fragmentation.
View Premium accesscourse lessons, concepts and answers · see current terms
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS
AI & GPU InfrastructureDesign a GPU serving platform for several LLMs with autoscaling and a cost ceiling.→Linux, Networking & ScriptingA Linux host is slow and someone says it is out of memory. How do you check, and what do free and df actually tell you?→AI & GPU InfrastructureDevice plugin versus Dynamic Resource Allocation for GPUs: why did Kubernetes need DRA?→AI & GPU InfrastructureFour teams want GPUs and you have twelve A100s. Walk me through MIG, time-slicing and MPS, and how you would decide.→AI & GPU InfrastructureA multi-node training job sits Pending forever while the cluster shows free GPUs. What is happening?→AI & GPU InfrastructureOne node in your training fleet makes every job it touches 30 percent slower, but it passes health checks. How do you find and handle it?→