Skip to content
chiltepin

Generated from: “How the orders service is deployed on Kubernetes.”

Orders service on Kubernetes deployment diagram

AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.

Edited 2026-09-14: Changed the relay to an application polling loop and removed unsupported recovery-time guarantees.

DOCUMENTDRAFT

Orders service on Kubernetes

What runs where in the production cluster, how a new version reaches traffic, and how it comes back out.

This example places three workloads in one namespace on an EKS cluster. Durable state lives in managed services outside the cluster. Pod replacement and node drains still require graceful shutdown and retry handling. A new version reaches customers through a canary that Argo Rollouts drives from the same metrics the SLO alerts use.

SECTION 01 · Note

Assumptions

Note
The request named the service and the platform, nothing else. This document assumes EKS 1.31 in us-east-1, ingress-nginx as the entry, Argo Rollouts for progressive delivery, and Vault Agent for secrets. Postgres is RDS, Redis is ElastiCache, Kafka is MSK. Replica counts are the steady-state values on 12 Sep 2026; the HPA changes them during the day.

What runs in the cluster

Only stateless workloads run in the cluster. orders-api serves HTTP; orders-worker consumes Kafka; outbox-relay is a long-running Deployment whose application loop polls committed outbox rows every 10 seconds. It marks rows delivered after a successful publish; consumers must tolerate duplicate delivery. The Vault node represents injected credentials rather than a fourth workload. Dashed edges mark asynchronous communication.

The polling interval belongs to the application. A Kubernetes CronJob schedule uses minute-level fields and cannot directly express an every-ten-seconds schedule.

SECTION 02 · Architecture

orders namespace on prod-use1

K8S
Block diagram: 8 nodes, 8 connectionsEKS · prod-use1AWS managed, outside the clusterorders namespaceingress-nginx namespaceingress-nginxNLB in frontINGRESSorders-apiGo · Argo RolloutDEPLOY×12orders-workerGo · DeploymentDEPLOY×4outbox-relayapplication pollsevery 10 sDEPLOYVault Agentsidecar in everypodSECRETSRDS Postgres16 · Multi-AZDBElastiCache Redis7 · 2 shardsCACHEMSK Kafka3 brokersBUS12345678
1HTTPS, host orders.example.com2reads and writes orders3caches carts, 15 min TTL4reads the outbox table5publishes orders.* events6consumes orders.*7updates fulfilment status8injects DB and Kafka credentials
LegendINGRESSingressDEPLOYdeploymentSECRETSsecretsDBdatabaseCACHEcacheBUSstream×Nreplicascallsasync / optionalentry point

How a version reaches traffic

The proposed rollout uses analysis gates for canary 5xx share and p99 latency. Configure and test the analysis templates and traffic routing before relying on automatic abort. Recovery time depends on that configuration and controller reconciliation.

SECTION 03 · Rollout

orders-api canary

ROLLOUTcanary
Stage 1
Smokedone
5%
10m
Stage 2
Canarycurrent
25%
20m
Stage 3
Halfnext
50%
20m
Stage 4
Fullnext
100%
LegenddonecurrentnextGATEgate — must pass to advance
RollbackRun kubectl argo rollouts abort orders-api -n orders, then verify traffic and health on the stable revision. Pod retention and recovery time depend on the rollout configuration.

The numbers each pod runs with

These resource settings are illustrative starting values. Choose requests, limits, and replica counts from load tests for the actual service. The disruption budget constrains voluntary evictions; it is not a guarantee of capacity during a zone outage.

SECTION 04 · Spec

orders-api runtime settings

Requests
500m CPU, 512 Mi memory per pod
Limits
1 CPU, 1 Gi memory; monitor throttling and out-of-memory terminations
HPA
min 12, max 120, target 60% CPU; scale-up stabilization 0 s, scale-down 5 min
PodDisruptionBudget
minAvailable 9
Topology spread
maxSkew 1 across the three zones; DoNotSchedule
Probes
readiness GET /healthz/ready every 5 s; liveness GET /healthz/live every 10 s, 3 failures restart
Shutdown
preStop sleeps 10 s so the ingress stops routing, then SIGTERM; terminationGracePeriod 30 s
Image
ghcr.io/example/orders-api, tag is the git SHA; no :latest anywhere

Rolling back by hand

The manual path is for failures the configured analysis does not catch. These commands use the Argo Rollouts kubectl plugin. Check the selected cluster and namespace before using them.

SECTION 05 · Steps

Roll orders-api back to the previous version

  1. Abort the current rollout

    Request an abort. Verify that traffic returns to the stable revision; retention depends on configuration.

    bash
    kubectl argo rollouts abort orders-api -n orders
  2. Confirm the stable version serves

    The stable revision must show 100% weight and the desired replica count.

    bash
    kubectl argo rollouts get rollout orders-api -n orders --watch

    Stop here if the customer report no longer reproduces. The next deploy starts a new canary from the fixed commit.

  3. Undo to the previous revision if stable is also bad

    This promotes the revision before the current stable, with the same canary gates.

    bash
    kubectl argo rollouts undo orders-api -n orders --to-revision=<n>
  4. Tell the channel

    Post the rollout name, the two revisions, and the reason in

View the Markdown
```meta
title: Orders service on Kubernetes
subtitle: What runs where in the production cluster, how a new version reaches traffic, and how it comes back out.
tag: DRAFT
```

This example places three workloads in one namespace on an EKS cluster. Durable state lives in managed services outside the cluster. Pod replacement and node drains still require graceful shutdown and retry handling. A new version reaches customers through a canary that Argo Rollouts drives from the same metrics the SLO alerts use.

```callout
tone: note
title: Assumptions
body: "The request named the service and the platform, nothing else. This document assumes EKS 1.31 in us-east-1, ingress-nginx as the entry, Argo Rollouts for progressive delivery, and Vault Agent for secrets. Postgres is RDS, Redis is ElastiCache, Kafka is MSK. Replica counts are the steady-state values on 12 Sep 2026; the HPA changes them during the day."
```

## What runs in the cluster

Only stateless workloads run in the cluster. `orders-api` serves HTTP; `orders-worker` consumes Kafka; `outbox-relay` is a long-running Deployment whose application loop polls committed outbox rows every 10 seconds. It marks rows delivered after a successful publish; consumers must tolerate duplicate delivery. The Vault node represents injected credentials rather than a fourth workload. Dashed edges mark asynchronous communication.

The polling interval belongs to the application. A Kubernetes [CronJob schedule](https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/#schedule-syntax) uses minute-level fields and cannot directly express an every-ten-seconds schedule.

```block
id: orders-k8s
preset: k8s
title: orders namespace on prod-use1
groups:
  - { id: cluster, col: 1, row: 1, cols: 3, rows: 3, label: "EKS · prod-use1" }
  - { id: ns-ingress, parent: cluster, col: 1, row: 1, cols: 1, rows: 3, label: "ingress-nginx namespace" }
  - { id: ns-orders, parent: cluster, col: 2, row: 1, cols: 2, rows: 3, label: "orders namespace" }
  - { id: aws, col: 4, row: 1, cols: 1, rows: 3, label: "AWS managed, outside the cluster" }
nodes:
  - { id: ingress, col: 1, row: 2, kind: ingress, name: ingress-nginx, tech: "NLB in front" }
  - { id: api, col: 2, row: 1, kind: deployment, name: orders-api, tech: "Go · Argo Rollout", replicas: 12 }
  - { id: worker, col: 2, row: 2, kind: deployment, name: orders-worker, tech: "Go · Deployment", replicas: 4 }
  - { id: relay, col: 2, row: 3, kind: deployment, name: outbox-relay, tech: "application polls every 10 s" }
  - { id: vault, col: 3, row: 2, kind: secret, name: Vault Agent, tech: "sidecar in every pod" }
  - { id: pg, col: 4, row: 1, kind: postgres, name: RDS Postgres, tech: "16 · Multi-AZ" }
  - { id: redis, col: 4, row: 2, kind: redis, name: ElastiCache Redis, tech: "7 · 2 shards" }
  - { id: kafka, col: 4, row: 3, kind: kafka, name: MSK Kafka, tech: "3 brokers" }
edges:
  - ingress -> api: "HTTPS, host orders.example.com"
  - api -> pg: "reads and writes orders"
  - api -> redis: "caches carts, 15 min TTL"
  - relay -> pg: "reads the outbox table"
  - relay --> kafka: "publishes orders.* events"
  - kafka --> worker: "consumes orders.*"
  - worker -> pg: "updates fulfilment status"
  - vault --> api: "injects DB and Kafka credentials"
```

## How a version reaches traffic

The proposed rollout uses analysis gates for canary 5xx share and p99 latency. Configure and test the analysis templates and traffic routing before relying on automatic abort. Recovery time depends on that configuration and controller reconciliation.

```rollout
id: orders-api-rollout
title: orders-api canary
strategy: canary
stages:
  - "[done] 5% · Smoke · 10m — no 5xx on the canary; readiness passes on every pod"
  - "[current] 25% · Canary · 20m — 5xx share < 0.5% and p99 < 800 ms on canary pods"
  - "[next] 50% · Half · 20m — same gates; error budget burn under 2× on the whole service"
  - "[next] 100% · Full — stable ReplicaSet scales to zero after 10 minutes"
rollback: "Run kubectl argo rollouts abort orders-api -n orders, then verify traffic and health on the stable revision. Pod retention and recovery time depend on the rollout configuration."
```

## The numbers each pod runs with

These resource settings are illustrative starting values. Choose requests, limits, and replica counts from load tests for the actual service. The disruption budget constrains voluntary evictions; it is not a guarantee of capacity during a zone outage.

```spec
id: orders-api-resources
title: orders-api runtime settings
accent: green
rows:
  - { label: Requests, value: "500m CPU, 512 Mi memory per pod" }
  - { label: Limits, value: "1 CPU, 1 Gi memory; monitor throttling and out-of-memory terminations" }
  - { label: HPA, value: "min 12, max 120, target 60% CPU; scale-up stabilization 0 s, scale-down 5 min" }
  - { label: PodDisruptionBudget, value: "minAvailable 9" }
  - { label: Topology spread, value: "maxSkew 1 across the three zones; DoNotSchedule" }
  - { label: Probes, value: "readiness GET /healthz/ready every 5 s; liveness GET /healthz/live every 10 s, 3 failures restart" }
  - { label: Shutdown, value: "preStop sleeps 10 s so the ingress stops routing, then SIGTERM; terminationGracePeriod 30 s" }
  - { label: Image, value: "ghcr.io/example/orders-api, tag is the git SHA; no :latest anywhere" }
```

## Rolling back by hand

The manual path is for failures the configured analysis does not catch. These commands use the [Argo Rollouts kubectl plugin](https://argo-rollouts.readthedocs.io/en/stable/features/kubectl-plugin/). Check the selected cluster and namespace before using them.

```steps
id: orders-rollback
title: Roll orders-api back to the previous version
items:
  - title: Abort the current rollout
    body: Request an abort. Verify that traffic returns to the stable revision; retention depends on configuration.
    code: kubectl argo rollouts abort orders-api -n orders
    lang: bash
  - title: Confirm the stable version serves
    body: The stable revision must show 100% weight and the desired replica count.
    code: kubectl argo rollouts get rollout orders-api -n orders --watch
    lang: bash
    note: Stop here if the customer report no longer reproduces. The next deploy starts a new canary from the fixed commit.
  - title: Undo to the previous revision if stable is also bad
    body: This promotes the revision before the current stable, with the same canary gates.
    code: kubectl argo rollouts undo orders-api -n orders --to-revision=<n>
    lang: bash
  - title: Tell the channel
    body: Post the rollout name, the two revisions, and the reason in #orders-oncall. The postmortem links to that message.
```