Uber separates scaling intent from execution on its Kubernetes platform
Uber added a ServiceScale resource and a separate controller so failover and other orchestrators can request workload scaling without putting rare failover logic on the deployment controller’s normal hot path. The shared intent stays visible in Kubernetes rather than moving into a separate database or coordination service.
The production lessons are the useful part: informer caches can lag, and two controllers writing the same workload exposed a ReplicaSet metadata/spec inconsistency that could break proportional scaling or leave workloads stuck. Uber added a generation-based read-your-own-write guardrail, drift monitoring and an automated healer, then rolled the change out through staging and canaries. InfoQ has no article comment section.