Internal SLOs: what a platform owes the teams building on it
A platform capability that other teams depend on has an implied objective whether or not one is published. Publishing it in the consumer's units, and defining what the platform stops doing on breach, converts an assumption everyone makes differently into a contract teams can design against.
TL;DR: State the objective as something a team can do, not as a component being up, and attach a consequence that costs the platform team something. An internal SLO with no breach behaviour is a dashboard, because the usual enforcement mechanism between a provider and a customer does not exist inside one company.
The failure this concept exists to prevent
A payments team commits to 99.95 percent availability. Their deploys go through the platform's pipeline. Secrets come from its vault, images from its registry. Nobody has told them what any of those are expected to deliver.
So they assume. Usually they assume the platform is more reliable than their own service, because it is centrally run and looks solid from outside. Their 99.95 percent rests on that assumption, and nobody has tested it.
The vault has an unannounced two-hour maintenance window. The payments service cannot start new instances during it. The first anyone learns that the assumption was wrong is during an incident that the payments team then owns, explains and is measured on.
State it in the consumer's units
A component-level claim cannot be designed against. "The registry is available" leaves a team unable to answer the only question they have, which is whether their service can start.
Write the objective as a user action with a time bound. A team can deploy a merged change within 10 minutes, 99.5 percent of the time. A running pod can fetch a secret within 2 seconds, 99.9 percent of the time. Those fail correctly when every component is nominally up and the path between them is broken, which is the common case and the one component metrics miss.
The arithmetic sets what you can promise
A 30-day month is about 43,200 minutes, so the permitted unavailability is:
| Objective | Permitted per month | What it implies |
|---|---|---|
| 99.5% | 216 min | A person can respond, diagnose and fix |
| 99.9% | 43.2 min | A rota, runbooks, and fast rollback |
| 99.99% | 4.32 min | Automated failover; no human in the loop |
The step from 99.9 to 99.99 is not a tighter target, it is a different system and a different headcount. A platform team of four cannot deliver four minutes a month on a capability with a single control plane, and writing the number down does not change that.
This table says nothing about which row you need. That comes from the other direction: ask the teams with the strictest commitments what they can absorb, and derive the platform objective from the obligations already made above it.
The consequence is the part that gets skipped
A cloud vendor that misses its objective pays a credit. The credit is small and nobody buys for it, but it establishes that the promise has a price. Inside a company there is no invoice, so the enforcement has to be behavioural and agreed before the breach.
Three that work, because each costs the platform team something real. Feature work pauses until a full window passes cleanly. A written review goes to every affected team, not just the one that noticed. And teams scheduled to migrate onto the breaching capability may defer, which is the one that teams actually believe, because it slows down the platform's own roadmap.
Self-check
Your platform publishes "deploy within 10 minutes of merge, 99.5 percent monthly". This month there were 4,000 deploys, of which 38 took longer than 10 minutes, and 22 of those 38 were during one 50-minute incident caused by a single team's runaway build consuming the shared runners. Decide whether you breached, and what you do about the 22. Work both through before reading on.
Start with the measurement. 38 / 4,000 = 0.95 percent of deploys missed the target, against a permitted 0.5 percent, so the objective is breached by roughly a factor of two. The permitted count was 4,000 x 0.005 = 20 and you spent 38.
The tempting move is to exclude the 22, on the grounds that another team caused them. Resist it. The objective is a promise about what a team experiences, and a team whose deploy took 25 minutes had that experience regardless of which neighbour caused it. An objective with cause-based carve-outs stops predicting anything, because every incident has a cause that can be described as somebody else's.
So you breached, and the breach consequences apply. What the 22 do change is the remedy. Sixteen scattered failures suggest a capacity or reliability problem in the pipeline itself. Twenty two from one noisy neighbour is an isolation problem, and the fix is per-team concurrency limits or dedicated runner pools rather than more capacity.
What the numbers do not settle is whether 99.5 was the right objective. If this is the first breach in a year it was about right and you have found a specific isolation defect. If it is the fourth, the platform has been promising something its architecture does not support, and the honest correction is to republish a lower number alongside the plan to earn the higher one back.
Next: error budgets and burn rate covers the arithmetic of spending an allowance, which applies to a platform objective exactly as it does to a service.