What breaks deploys
The 140 failed deploys of 2026 by root cause, ranked so the few causes behind most failures stand out.
The 140 failed deploys this year come from eight root causes, but three of them account for 99 failures. Fixing those three removes 71% of the failures. The other five causes together fail fewer deploys than the second cause alone.
Assumptions
checkout-api to the prod cluster between 1 January and 12 September 2026. A deploy failed when the Argo CD rollout aborted or the post-deploy smoke test failed. Each failure carries one root cause, taken from the incident label that the on-call engineer set when the pipeline was rerun. Failures without a label were re-labeled from the pipeline log; 5 of them did not fit any category and are counted as Other.The headline
Config drift, failed migrations, and flaky integration tests. The remaining five causes explain 41.
Failures by root cause
The bars are sorted by count and the line is the running share. The line crosses 80% at the fourth cause, so anything past expired secrets is a long tail and not worth a project of its own.
Failed deploys by root cause, 2026 to date
What each cause costs and what fixes it
Count is not the only measure. A failed migration takes about three times longer to recover than a flaky test, because the rollback needs a human to check the schema first. The fix column names the change that removes the cause, not a mitigation.
Root causes with recovery time and fix
| Root cause | Failures | Share | Cumulative | Median recovery | Fix that removes the cause |
|---|---|---|---|---|---|
| Config drift between staging and prod | 46 | 33% | 33% | 25 min | Generate both environments from one values file with a diff gate in CI |
| Failed database migration | 31 | 22% | 55% | 70 min | Run every migration against a nightly prod snapshot before merge |
| Flaky integration test | 22 | 16% | 71% | 20 min | Quarantine tests that fail and pass on rerun; retry once inside the gate |
| Expired secret or certificate | 14 | 10% | 81% | 40 min | Alert 14 days before expiry; read secrets from the vault at start-up |
| Image pull timeout | 9 | 6% | 87% | 15 min | Pre-pull the image to every node before the rollout starts |
| Runner capacity timeout | 7 | 5% | 92% | 65 min | Add a dedicated runner pool for deploy jobs |
| Canary latency abort | 6 | 4% | 96% | 30 min | Keep. Four of the six aborts were correct |
| Other | 5 | 4% | 100% | 35 min | No shared fix |
Shares are rounded to whole percent, so the cumulative column can differ from the sum of the shares by one point. Recovery is the median time from the failed run to the next successful rollout of the same commit.
What to do with this
- 1Fix config drift first
It causes one failure in three and the fix is a CI gate, not a platform change.
- 2Failed migrations are the second cause and the slowest to recover
31 failures at a 70 minute median is more lost time than config drift.
- 3Three fixes remove 71% of failures
The pareto chart at
- 4Leave the canary gate alone
Its aborts are mostly correct. Removing them would trade six failed deploys for four bad releases.