54Half your inference fleet went unschedulable overnight and nobody deployed anything. Start.▼hardNewNVIDIAMicrosoftOracle2 replies◆ premiumThe pods are Pending, the nodes are Ready, and the GPUs have vanished from allocatable. Something changed on the node under a fleet that nobody considered part of the deploy surface.Open full answer →
05You need to upgrade GPU drivers across a live inference and training fleet. What is your plan?▼hardNewNVIDIAGoogleMicrosoft2 repliesunlockedThe driver, the container toolkit, the CUDA runtime inside the image and the framework build all have to agree. Upgrade the wrong one first and every pod on the node fails to start with an error that names none of them.Open full answer →