← DevOps
Technique Deployment and runtime lifecycle

Startup, readiness,
and liveness probes

A failed check is a request to do something.

An incident report links ticket request timeouts, a slow database, API replica restarts, and a rise in reconnect work. That sequence may mean a health check turned a dependency problem into restart pressure. It may also hide a process exit or node event. Follow the evidence before changing what the controller is allowed to do.

The judgment to keep

Decide whether a failure calls for more startup time, less incoming traffic, or a fresh process. Then choose a signal that justifies that action.

Kubernetes · one service · three failure drills
01 / Read the incident

A restart pattern is a clue. Trace the check that asked for it.

The incident report says the ticket database slowed down, API replicas restarted, and reconnect work rose during recovery. That sequence deserves attention, but it does not yet prove the probe caused the restarts. A process exit, memory kill, node disruption, or rollout can happen near the same dependency incident.

A process can be running while it initializes, running while a database is unavailable, or running while its work is stuck. A single green/red flag collapses those states. You need to decide which response would help before writing the check.

A probe is a repeated check whose result feeds a controller’s decision. In Kubernetes, the kubelet runs container probes. A failing result is therefore part of the control loop, not just a status message on a dashboard.

Diagnose before tuning: list another plausible cause for the restarts. Which record would tell you whether the kubelet restarted a container after liveness failures, or the process exited for another reason?
Compare a first diagnostic passReveal after making a prediction
Diagnostic checkpoint Keep the event sequence separate from its explanation.
Possible causes
A dependency-backed liveness check failed; the process crashed or was killed; or a rollout/node event replaced it.
Discriminating evidence
Record the initial restart count, inspect Pod events and the previous container's termination reason, then compare timestamps with the configured probe path and dependency errors.
What this fixture establishes
The intentionally flawed probe example points liveness at /readyz, whose contract includes database availability. That configuration predicts the failure path. It is not evidence that any live cluster followed it; the disposable drill must supply those observations.
Three questions, three consequences
ProbeQuestionWhat failure asks Kubernetes to do
StartupHas this container finished starting?Keep readiness and liveness gated. If the startup failure threshold is reached, terminate the container; the restart policy governs restart.
ReadinessShould this Pod receive new Service traffic?Mark it unready after the configured failures. Ordinary Service routing excludes unready endpoints; the container keeps running.
LivenessHas this container entered a state where restarting is warranted?Terminate it after the configured failures. With a Deployment’s restart policy, it restarts.

Once startup succeeds, its check stops for that container lifetime. Readiness can change repeatedly. Liveness runs independently of readiness after the startup gate opens; an unready container can still fail liveness and restart. These are the platform’s distinct probe semantics.

Operating case / Ticket APIThree replicas, one required database.
User promise
A ticket request needs a working database; there is no useful cached response in this example.
Startup
Local configuration and an in-memory index take up to 50 seconds to initialize in the fixture.
Recoverable
The process can reconnect when a temporary database outage ends.
Restartable
A stuck local work loop stops progress and recovers in a fresh process.
Leave with: the action you want for each failure. A dependency outage and a deadlock should not accidentally request the same action.
02 / Define the signals

An endpoint should answer the question its caller is asking.

Give this service three explicit endpoint contracts. Return a small 200 response for success and 503 for failure. Keep the endpoints cheap, unauthenticated for the kubelet’s access path, and free of sensitive diagnostics. Avoid redirecting a probe to a login page.

The application implements these meanings; endpoint names alone do nothing
EndpointSuccess meansWork to avoid
/startupzLocal initialization has completed. This remains true until this process exits.Making completion depend on every remote service being available.
/readyzThe service is initialized, local work can progress, and its required database is usable.A full user transaction or an unbounded dependency call on every probe.
/livezThe local execution path being monitored can still make progress.Database, DNS, or third-party checks that fail in every replica together.

In a real service, readiness might read a recent, bounded dependency observation from a background monitor. Define how old that observation may be and how recovery refreshes it. A cached “good” result that never expires is a false promise; synchronous probes that pile up slow database calls can become extra load during the outage.

For liveness, choose evidence tied to useful local progress. An HTTP response from the request event loop can reveal that loop being blocked; a separate health thread returning 200 says little about a deadlocked worker. If a heartbeat drives the check, define when progress is expected so an idle worker is not mistaken for a stuck one.

Leave with: a sentence for each endpoint that another engineer could turn into a test.
03 / Set the tolerances

Give slow startup its own allowance.

The fixture takes 50 seconds to start. Give startup about 120 seconds of nominal allowance: 24 failed samples at a five-second period. That leaves room for a slow cold start without forcing the running service’s liveness check to wait two minutes before noticing a fault.

This excerpt lives under the container in a Deployment. The named port http maps to port 8080. The complete manifest and disposable service appear in the drill below.

deployment.yaml · container probe configuration
startupProbe:
  httpGet:
    path: /startupz
    port: http
  periodSeconds: 5
  timeoutSeconds: 1
  failureThreshold: 24
readinessProbe:
  httpGet:
    path: /readyz
    port: http
  periodSeconds: 5
  timeoutSeconds: 1
  failureThreshold: 2
  successThreshold: 1
livenessProbe:
  httpGet:
    path: /livez
    port: http
  periodSeconds: 5
  timeoutSeconds: 1
  failureThreshold: 3

timeoutSeconds bounds one attempt. failureThreshold counts consecutive failed attempts before action. A success breaks that failure streak. Readiness also uses successThreshold to control recovery; it is one here. Startup and liveness require a success threshold of one. See the Kubernetes configuration reference.

The two readiness failures tolerate a brief blip before withdrawing traffic. The three liveness failures tolerate a little more before restart. Their nominal windows are about 10 and 15 seconds; they are not precise deadlines measured from the onset of a fault. Probe scheduling, timeouts, and the phase of the first failure affect detection. Termination and endpoint propagation add their own delays.

Choose these numbers from cold-start measurements, normal local pauses, probe latency, and the time users can tolerate a broken replica. CPU throttling or overload can delay a healthy process. Aggressive thresholds can then remove capacity at exactly the wrong time.

Why not just add a long initial delay?Startup gating and steady operation

A fixed delay protects startup only for that fixed duration. A startup probe can succeed early when initialization is quick, while still allowing a bounded slow start. It then gets out of the way. If initialization never succeeds, investigate it instead of increasing the allowance indefinitely.

A readiness check alone does not suppress liveness. If you configure liveness to fail while initialization is still legitimate, the process can be restarted before it ever finishes.

Leave with: tolerances tied to observed service behavior, and a clear distinction between an approximate detection window and recovery time.
04 / Follow the outage

A restart can add a second problem to the first.

Reusing /readyz for liveness looks tidy. For this service, it turns database failure into a request to restart the application. Nothing in that action makes the database available again.

Counterexample · database outage becomes container restarts
livenessProbe:
  httpGet:
    path: /readyz # Wrong for this service: this includes the database.
    port: http
  periodSeconds: 5
  timeoutSeconds: 1
  failureThreshold: 3
01 / TriggerDatabase stalls

Ticket requests cannot complete.

02 / MisdiagnosisLiveness fails

Every replica can see the same remote fault.

03 / ActionContainers restart

Local initialization and connection setup repeat.

04 / FeedbackRecovery gets harder

Connection churn and lost warm capacity may prolong the outage.

That last step is a risk to investigate, not an automatic numerical claim. This page’s experiment counts restarts in a small model; it does not measure database load or outage duration. It makes the unnecessary intervention visible.

Follow the consequence

One failure, two restart policies.

Database unavailable from 20 s until 100 s. Both policies start with a ready replica.

Local liveness

Traffic withdrawn
Container
running
Restarts so far
0
Liveness failures
0 / 3
Startup failures
0 / 24

Database unavailable; leave the recoverable process alive.

Restarts over the full 120 s: 0

Liveness also checks the database

Traffic withdrawn
Container
starting
Restarts so far
1
Liveness failures
0 / 3
Startup failures
0 / 24

Restart requested. Initialization begins again; traffic stays off.

Restarts over the full 120 s: 2

Authored, tested model: aligned 5 s samples, two readiness failures, three liveness failures. Restart is immediate; warmup after restart is 20 s (50 s in the startup case). Kubernetes adds scheduling, termination, backoff, and routing delays. These timestamps are model outputs, not a cluster recording.

In the database scenario, local liveness leaves the process running. Readiness withdraws traffic, and the same process can become ready when the dependency returns. The database-dependent policy restarts twice in the authored schedule. In the stuck-process scenario, a restart is useful: both policies recover local progress.

If all three replicas depend on the same failed database, all three may become unready. There is still a user-visible outage. Readiness does not manufacture a working replica; keeping liveness local prevents additional damage while you repair the dependency or activate a designed fallback.

Leave with: a causal explanation of what restarting changes, what it leaves broken, and what extra work it creates.
05 / Exercise the failures

Check the controller’s response, not just the HTTP status.

A health handler test can prove which status code a condition produces. It cannot prove that your manifest points to that handler, the kubelet reaches the port, or traffic stops going to an unready endpoint. Verify the application contract and the deployed reaction separately.

Failure drill and acceptance evidence
IntroduceObserveExpected consequence
50-second startupStartup results, Pod readiness, restart countUnready during initialization; no premature restart; ready after initialization and readiness success.
Database unavailableReadiness failures, EndpointSlice readiness conditions, restart countAffected endpoints become unready; restart count stays unchanged.
Database restoredReadiness recovery and a ticket request through the ServiceTraffic returns without replacing the container.
Local progress stuckLiveness failure events and container restart countThe affected container restarts, passes startup again, and returns to readiness.

Capture the initial restart count so an earlier crash does not look like a new liveness action. Inspect Pod events and the previous container’s termination reason. An out-of-memory kill, application exit, and liveness-triggered restart need different fixes.

Endpoint readiness is control-plane evidence. Also send representative requests through the Service from another Pod and watch which replica serves them. Direct Pod access and port forwarding bypass ordinary Service selection; they cannot establish that unready replicas have stopped receiving Service traffic.

Run a disposable local drillComplete fixture · Docker, kind, and kubectl

The files live in src/lib/content/lessons/startup-readiness-and-liveness-probes/examples. Work through the commands in that directory one observation at a time. They create an explicitly named local kind cluster and namespace, load the fixture image, and provide cleanup.

The fixture uses marker files to simulate database failure and a detected local stall. It does not connect to a real database or deliberately deadlock the event loop. Endpoint behavior and the browser model have automated tests; a live cluster drill remains a separate check.

drill.sh · run one observation at a time
# Run from this lesson's examples directory. Requires Docker, kind, and kubectl.
# All cluster commands explicitly target a disposable local cluster.
kind create cluster --name probes-lesson
docker build -t probes-lesson:v1 .
kind load docker-image probes-lesson:v1 --name probes-lesson
kubectl --context kind-probes-lesson create namespace probes-lesson
kubectl --context kind-probes-lesson apply -f deployment.yaml
kubectl --context kind-probes-lesson -n probes-lesson rollout status deployment/tickets --timeout=180s

# Observe before injecting a failure. Record restart counts for each Pod.
kubectl --context kind-probes-lesson -n probes-lesson get pods -l app=tickets
kubectl --context kind-probes-lesson -n probes-lesson get endpointslices -l kubernetes.io/service-name=tickets -o yaml

# Start an in-cluster client. Its image is already loaded; this adds no new image.
kubectl --context kind-probes-lesson -n probes-lesson run request-check --image=probes-lesson:v1 --image-pull-policy=Never --restart=Never --command -- node -e 'setInterval(() => {}, 1000)'
kubectl --context kind-probes-lesson -n probes-lesson wait --for=condition=Ready pod/request-check --timeout=60s
# Repeat this request before, during, and after the drill. Record status and replica.
kubectl --context kind-probes-lesson -n probes-lesson exec request-check -- node -e "fetch('http://tickets/tickets').then(async r => console.log(r.status, r.headers.get('x-replica'), await r.text()))"

# A marker simulates an unavailable dependency for ONE replica.
kubectl --context kind-probes-lesson -n probes-lesson exec deployment/tickets -- node -e "require('node:fs').writeFileSync('/tmp/database-down', '')"
# Wait at least 15 seconds, then inspect readiness, restarts, and endpoint conditions.
# No restart is expected. Save the affected Pod name from this output.
kubectl --context kind-probes-lesson -n probes-lesson get pods -l app=tickets
kubectl --context kind-probes-lesson -n probes-lesson get endpointslices -l kubernetes.io/service-name=tickets -o yaml

# Replace AFFECTED_POD with that name; do not let exec choose a different replica.
kubectl --context kind-probes-lesson -n probes-lesson exec AFFECTED_POD -- node -e "require('node:fs').unlinkSync('/tmp/database-down')"
# After recovery, simulate a local progress failure in that same replica.
kubectl --context kind-probes-lesson -n probes-lesson exec AFFECTED_POD -- node -e "require('node:fs').writeFileSync('/tmp/stuck', '')"
kubectl --context kind-probes-lesson -n probes-lesson describe pod AFFECTED_POD
# Expect a liveness failure, one container restart, startup gating, then Ready.
# Marker files live in the container writable layer and disappear on restart.

# Cleanup only this disposable cluster when observations are complete.
kind delete cluster --name probes-lesson
Complete Deployment and Service
deployment.yaml · complete local fixture
apiVersion: apps/v1
kind: Deployment
metadata:
  name: tickets
  namespace: probes-lesson
spec:
  replicas: 3
  selector:
    matchLabels:
      app: tickets
  template:
    metadata:
      labels:
        app: tickets
    spec:
      containers:
        - name: tickets
          image: probes-lesson:v1 # Build and load the local fixture first.
          imagePullPolicy: Never
          ports:
            - name: http
              containerPort: 8080
          env:
            - name: STARTUP_MS
              value: '50000'
          resources:
            requests:
              cpu: 100m
              memory: 64Mi
            limits:
              memory: 128Mi
          startupProbe:
            httpGet:
              path: /startupz
              port: http
            periodSeconds: 5
            timeoutSeconds: 1
            failureThreshold: 24
          readinessProbe:
            httpGet:
              path: /readyz
              port: http
            periodSeconds: 5
            timeoutSeconds: 1
            failureThreshold: 2
            successThreshold: 1
          livenessProbe:
            httpGet:
              path: /livez
              port: http
            periodSeconds: 5
            timeoutSeconds: 1
            failureThreshold: 3
---
apiVersion: v1
kind: Service
metadata:
  name: tickets
  namespace: probes-lesson
spec:
  selector:
    app: tickets
  ports:
    - port: 80
      targetPort: http
Health endpoint fixture
app.mjs · simulated health observations
// Disposable drill fixture. Marker files simulate observations, not a real DB or deadlock.
import { createServer } from 'node:http';
import { existsSync } from 'node:fs';
import { pathToFileURL } from 'node:url';

/**
 * @param {string | undefined} path
 * @param {{ initialized: boolean, stuck: boolean, databaseDown: boolean }} state
 */
export function statusFor(path, { initialized, stuck, databaseDown }) {
	if (path === '/startupz') return initialized ? 200 : 503;
	if (path === '/livez') return stuck ? 503 : 200;
	if (path === '/readyz' || path === '/tickets') {
		return initialized && !stuck && !databaseDown ? 200 : 503;
	}
	return 404;
}

if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) {
	const startedAt = Date.now();
	createServer((request, response) => {
		const status = statusFor(request.url, {
			initialized: Date.now() - startedAt >= Number(process.env.STARTUP_MS ?? 50000),
			stuck: existsSync('/tmp/stuck'),
			databaseDown: existsSync('/tmp/database-down')
		});
		response.writeHead(status, {
			'Content-Type': 'text/plain',
			'Cache-Control': 'no-store',
			'X-Replica': process.env.HOSTNAME ?? 'local'
		});
		response.end(status === 200 ? 'ok\n' : 'unavailable\n');
	}).listen(Number(process.env.PORT ?? 8080), '0.0.0.0');
}
Container image
Dockerfile · local learning image
# Learning fixture; pin an approved digest for your own deployment.
FROM node:22-alpine
WORKDIR /app
COPY app.mjs .
USER node
CMD ["node", "app.mjs"]

If the expected withdrawal or restart never happens, check the configured path, named port, probe events, and actual response first. If the response is correct but traffic persists, inspect the routing path and existing connections. If restarts happen without liveness failures, follow the termination reason instead of tuning probe thresholds.

Leave with: evidence linking the induced failure to readiness, endpoint selection, and restart behavior, including recovery.
06 / Leave an operating rule

Ask what a fresh process could actually repair.

Try a decision

The database is unavailable. What should these replicas report?

Ticket requests require the database. The application can reconnect without restarting, and its local work loop is still responsive.

Your next move

A liveness probe earns its place when it can detect a failure that a restart is likely to repair. If your process already exits reliably on unrecoverable local failure, the restart policy may be enough. Adding a second failure detector is a design decision with false positives and operating costs.

Keep monitoring alongside probes. Alert on user-visible errors and latency, sustained loss of ready capacity, and repeated restarts. A green health endpoint does not establish that real requests meet their promise.

Leave a short note with the configuration: what each endpoint means, where its signal comes from, why its tolerances fit, and how to exercise its failure and recovery. That note helps the next person resist “simplifying” three different questions into one.

The question to keep: if this check fails across every replica, does the action help the system recover?

Connections to follow nextRelated lessons

Retry, backoff, and idempotency explains why repeated attempts need a budget. Restarts can create a similar surge of repeated connection work.

Backpressure and queues examines what to do when incoming work exceeds capacity. Restarting overloaded workers may remove capacity instead of relieving the bottleneck.

Timeouts, deadlines, and races separates one attempt’s time limit from the whole operation’s budget.

Sources and scopePrimary documentation · checked September 27, 2026