DevOpsInterviewPrep logo
← 🧰 Platform & Cloud Economics
Foundational

Migration cost asymmetry: why a better platform does not get adopted

Teams decline to migrate to a better platform because the cost is immediate, certain and theirs while the benefit is deferred, uncertain and partly captured by others. The remedy is to move the cost rather than to repeat the argument.

TL;DR: A team deferring a migration every quarter is answering the question correctly each time, because the cost falls inside this quarter's committed work and the benefit does not. Platforms that get adopted move the cost onto themselves. Platforms that do not, send another announcement.

The failure this concept exists to prevent

The platform team ships something measurably better. Build times halve in the pilot. They present the numbers, write documentation, run a workshop, and six months later adoption sits at 11 percent.

The conclusion drawn in the retrospective is that teams are conservative or that communication was inadequate, so the next quarter brings a better announcement and a slide with a deadline. Adoption moves to 14 percent.

Nobody in this story behaved irrationally, and that is the part the retrospective misses.

The asymmetry, stated plainly

Three properties separate the cost from the benefit, and each pushes the same way.

Timing. The cost is paid this quarter, inside work already committed to a roadmap somebody else is tracking. The benefit accrues over following quarters.

Certainty. Three engineer-days is a number a team can predict. "Deploys will be faster" is a forecast made by the group that wants the migration to happen.

Incidence. The engineer who spends the three days bears the cost. The benefit is shared across the team, the platform's reported metrics and the organisation, and some of it is captured by people who paid nothing.

A rational actor facing an immediate certain private cost and a deferred uncertain shared benefit defers. The decision recurs every planning cycle and gets the same answer, which is why the service that is hardest to migrate is still on the old path three years later.

Moving the cost, with the numbers

The only reliable interventions change one of the three properties, and the strongest changes incidence by moving the work to the platform team.

Consider 80 services at an estimated 3 days each:

80 x 3 = 240 team-days if every team migrates itself

If 70 percent of the change is mechanical, the platform generates those and the teams review them:

80 x 0.7 = 56 services handled centrally, leaving 24 x 3 = 72 team-days distributed

The per-team ask falls from three days to roughly an hour of review for most services, which is small enough to fit in a sprint without a planning conversation. That is the entire mechanism: not persuasion, but making the ask smaller than the threshold at which teams have to negotiate.

What this does not account for is the platform team's cost of building the tooling, commonly two to four weeks. At 80 services that is clearly worth it. At 8 it is not, and embedding an engineer with each team is cheaper. The break-even is worth computing rather than assuming, because automating a migration is itself a project that can overrun.

Why deadlines land differently than they read

A deadline changes certainty, not incidence. The cost is still three days and still theirs, and the team now also carries the risk of missing a date. So the rational response shifts from deferring to doing it as late and as cheaply as possible, which produces exactly the badly-done compliance migrations that then become the platform team's support burden.

Deadlines work for the tail, after the economics have moved and most teams have gone. They are a tool for closing, not for opening.

Self-check

A platform team reports 11 percent adoption after six months and proposes a company-wide deadline in two quarters. One team migrated in week three and reports deploy times down from 40 minutes to 6. Say what the 11 percent and the one success together tell you, and decide what to do instead of the deadline. Work both through before reading on.

The 11 percent measures the asymmetry, not the platform. Taken with a successful migration that delivered a large and verifiable improvement, it says the product is good and the cost of getting to it is above the threshold teams will pay unprompted. Those two facts together point away from a communications problem and away from a quality problem.

The early adopter is the most useful asset in the story and is probably being used wrongly. A platform-authored case study reads as marketing. The same numbers from that team's own engineer, in their words, in a forum their peers already read, is evidence rather than advocacy, and it also tells you what the migration actually cost them, which is a number you need and do not have.

So the work is to measure the real per-service cost from that first migration, mechanise whatever share of it is repeatable, and go to the two or three teams with the worst current pain offering to do the work with them. Fix the economics first, and keep the deadline available for the last few services once most of the fleet has moved.

What none of this settles is whether the remaining 89 percent are all the same case. Some will be ordinary deferral, and some will be services with a real incompatibility that the pilot did not surface. Those are different problems, and the only way to tell them apart is to attempt a migration on a handful and find out.

Next: platform as a product covers how to read an adoption figure once you know what suppresses it.

RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS