28Explain the conntrack state machine. Which timers retire a connection, and why does the table fill with no leak?▼medium★ EssentialNewAmazon & AWSFlipkartRazorpay◆ premiumTable full is the symptom everyone knows. The interview answer lives one layer down: per-state timers, the five-day ESTABLISHED default, and what conntrack -S says about hash pressure.Open full answer →
47With tens of thousands of services, DNS lookups start timing out for five seconds at a stretch. Diagnose and fix Kubernetes DNS at scale.▼hard★ EssentialNewCloudflareUberLinkedIn◆ premiumThe classic five-second stall is rarely CoreDNS being slow; it is UDP, conntrack races and ndots expansion. The diagnosis ladder and the NodeLocal DNSCache fix separate operators from readers of architecture blogs.Open full answer →
06Requests inside the cluster fail with DNS errors, but only sometimes. Diagnose it.▼hardNewUberCloudflareShopify2 repliesunlockedIntermittent DNS is the most-reported and least-understood Kubernetes failure. There are three well-known causes and each leaves a different fingerprint.Open full answer →
14A busy node drops new connections intermittently; dmesg shows 'nf_conntrack: table full'. Walk me through what is happening.▼hard★ EssentialNewAmazon & AWSCloudflareUber○ sign inEstablished traffic keeps working while new connections die at random, which is exactly why this one confuses people. The table is full of flows that will not be needed again for days.Open full answer →