AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.
Edited 2026-09-14: Changed the relay to an application polling loop and removed unsupported recovery-time guarantees.
View the Markdown
```meta
title: Orders service on Kubernetes
subtitle: What runs where in the production cluster, how a new version reaches traffic, and how it comes back out.
tag: DRAFT
```
This example places three workloads in one namespace on an EKS cluster. Durable state lives in managed services outside the cluster. Pod replacement and node drains still require graceful shutdown and retry handling. A new version reaches customers through a canary that Argo Rollouts drives from the same metrics the SLO alerts use.
```callout
tone: note
title: Assumptions
body: "The request named the service and the platform, nothing else. This document assumes EKS 1.31 in us-east-1, ingress-nginx as the entry, Argo Rollouts for progressive delivery, and Vault Agent for secrets. Postgres is RDS, Redis is ElastiCache, Kafka is MSK. Replica counts are the steady-state values on 12 Sep 2026; the HPA changes them during the day."
```
## What runs in the cluster
Only stateless workloads run in the cluster. `orders-api` serves HTTP; `orders-worker` consumes Kafka; `outbox-relay` is a long-running Deployment whose application loop polls committed outbox rows every 10 seconds. It marks rows delivered after a successful publish; consumers must tolerate duplicate delivery. The Vault node represents injected credentials rather than a fourth workload. Dashed edges mark asynchronous communication.
The polling interval belongs to the application. A Kubernetes [CronJob schedule](https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/#schedule-syntax) uses minute-level fields and cannot directly express an every-ten-seconds schedule.
```block
id: orders-k8s
preset: k8s
title: orders namespace on prod-use1
groups:
- { id: cluster, col: 1, row: 1, cols: 3, rows: 3, label: "EKS · prod-use1" }
- { id: ns-ingress, parent: cluster, col: 1, row: 1, cols: 1, rows: 3, label: "ingress-nginx namespace" }
- { id: ns-orders, parent: cluster, col: 2, row: 1, cols: 2, rows: 3, label: "orders namespace" }
- { id: aws, col: 4, row: 1, cols: 1, rows: 3, label: "AWS managed, outside the cluster" }
nodes:
- { id: ingress, col: 1, row: 2, kind: ingress, name: ingress-nginx, tech: "NLB in front" }
- { id: api, col: 2, row: 1, kind: deployment, name: orders-api, tech: "Go · Argo Rollout", replicas: 12 }
- { id: worker, col: 2, row: 2, kind: deployment, name: orders-worker, tech: "Go · Deployment", replicas: 4 }
- { id: relay, col: 2, row: 3, kind: deployment, name: outbox-relay, tech: "application polls every 10 s" }
- { id: vault, col: 3, row: 2, kind: secret, name: Vault Agent, tech: "sidecar in every pod" }
- { id: pg, col: 4, row: 1, kind: postgres, name: RDS Postgres, tech: "16 · Multi-AZ" }
- { id: redis, col: 4, row: 2, kind: redis, name: ElastiCache Redis, tech: "7 · 2 shards" }
- { id: kafka, col: 4, row: 3, kind: kafka, name: MSK Kafka, tech: "3 brokers" }
edges:
- ingress -> api: "HTTPS, host orders.example.com"
- api -> pg: "reads and writes orders"
- api -> redis: "caches carts, 15 min TTL"
- relay -> pg: "reads the outbox table"
- relay --> kafka: "publishes orders.* events"
- kafka --> worker: "consumes orders.*"
- worker -> pg: "updates fulfilment status"
- vault --> api: "injects DB and Kafka credentials"
```
## How a version reaches traffic
The proposed rollout uses analysis gates for canary 5xx share and p99 latency. Configure and test the analysis templates and traffic routing before relying on automatic abort. Recovery time depends on that configuration and controller reconciliation.
```rollout
id: orders-api-rollout
title: orders-api canary
strategy: canary
stages:
- "[done] 5% · Smoke · 10m — no 5xx on the canary; readiness passes on every pod"
- "[current] 25% · Canary · 20m — 5xx share < 0.5% and p99 < 800 ms on canary pods"
- "[next] 50% · Half · 20m — same gates; error budget burn under 2× on the whole service"
- "[next] 100% · Full — stable ReplicaSet scales to zero after 10 minutes"
rollback: "Run kubectl argo rollouts abort orders-api -n orders, then verify traffic and health on the stable revision. Pod retention and recovery time depend on the rollout configuration."
```
## The numbers each pod runs with
These resource settings are illustrative starting values. Choose requests, limits, and replica counts from load tests for the actual service. The disruption budget constrains voluntary evictions; it is not a guarantee of capacity during a zone outage.
```spec
id: orders-api-resources
title: orders-api runtime settings
accent: green
rows:
- { label: Requests, value: "500m CPU, 512 Mi memory per pod" }
- { label: Limits, value: "1 CPU, 1 Gi memory; monitor throttling and out-of-memory terminations" }
- { label: HPA, value: "min 12, max 120, target 60% CPU; scale-up stabilization 0 s, scale-down 5 min" }
- { label: PodDisruptionBudget, value: "minAvailable 9" }
- { label: Topology spread, value: "maxSkew 1 across the three zones; DoNotSchedule" }
- { label: Probes, value: "readiness GET /healthz/ready every 5 s; liveness GET /healthz/live every 10 s, 3 failures restart" }
- { label: Shutdown, value: "preStop sleeps 10 s so the ingress stops routing, then SIGTERM; terminationGracePeriod 30 s" }
- { label: Image, value: "ghcr.io/example/orders-api, tag is the git SHA; no :latest anywhere" }
```
## Rolling back by hand
The manual path is for failures the configured analysis does not catch. These commands use the [Argo Rollouts kubectl plugin](https://argo-rollouts.readthedocs.io/en/stable/features/kubectl-plugin/). Check the selected cluster and namespace before using them.
```steps
id: orders-rollback
title: Roll orders-api back to the previous version
items:
- title: Abort the current rollout
body: Request an abort. Verify that traffic returns to the stable revision; retention depends on configuration.
code: kubectl argo rollouts abort orders-api -n orders
lang: bash
- title: Confirm the stable version serves
body: The stable revision must show 100% weight and the desired replica count.
code: kubectl argo rollouts get rollout orders-api -n orders --watch
lang: bash
note: Stop here if the customer report no longer reproduces. The next deploy starts a new canary from the fixed commit.
- title: Undo to the previous revision if stable is also bad
body: This promotes the revision before the current stable, with the same canary gates.
code: kubectl argo rollouts undo orders-api -n orders --to-revision=<n>
lang: bash
- title: Tell the channel
body: Post the rollout name, the two revisions, and the reason in #orders-oncall. The postmortem links to that message.
```