← DevOps
Technique Deployment and runtime lifecycle

Graceful shutdown
and connection draining

A process exit is part of the service contract.

During a routine rollout, a checkout API reports a small burst of interrupted requests. The new Pods become Ready and the Deployment completes, but that green result does not explain what happened to requests already in flight. Was the old process killed too quickly, did traffic keep reaching it, or did a client retry turn one interruption into two charges?

The judgment to keep

Separate new requests from work already accepted. Give each a bounded path to finish, then use request and termination evidence to decide whether the path worked.

Kubernetes · TypeScript · Go · rollout evidence
01 / Read the report

“The rollout was green” and “a few requests failed” can both be true.

Suppose support reports several checkout requests ending with connection resets during a rollout. The Deployment became available, error rates returned to baseline, and the team has no confirmed duplicate orders. This is a useful starting report; it is not yet evidence that the process exited too early or that Kubernetes sent traffic to a terminating Pod.

Keep the questions separate: which requests failed, which Pod handled them, whether the process was terminating, whether that request had already been accepted, and what the client did next. A retry may succeed while the first attempt is still completing, which makes idempotency part of the investigation.

Diagnose before tuning: what single timeline would you build from the client request ID, Pod name, application logs, termination events, and rollout timestamps?
Compare a first diagnostic passReveal after making a prediction
Diagnostic checkpointCorrelated timing narrows a search; it does not prove a cause.
Competing explanations
The process exited while work was active; new traffic reached a Pod during route convergence; a handler or dependency timed out; or the client abandoned and retried the request.
Useful evidence
Request IDs and durations, response/reset type, Pod UID, termination start and exit time, application signal logs, EndpointSlice conditions, and client retry records.
What the lesson establishes
The sample configuration defines a lifecycle to investigate. It does not claim a real cluster had these timings or that every client follows the same retry policy.
02 / Trace termination

Withdrawal and process shutdown overlap in time.

When a Pod is deleted, Kubernetes starts its termination grace countdown and begins removing the Pod from matching EndpointSlices while running any preStop hook. The hook runs synchronously; after it completes, the container runtime sends the stop signal (commonly SIGTERM). At grace expiry, remaining processes can be forcefully terminated.

EndpointSlice updates and application shutdown are concurrent parts of the transition. Endpoint changes need to propagate through consumers and proxies; they do not form a guarantee that all traffic has stopped before the process receives its signal. This is why the handler should stop accepting new work promptly and remain robust to requests that arrive during convergence.

“Drain” has multiple meanings. Stop accepting new requests on a listener; let accepted requests finish; close or wait for keep-alive connections; separately handle upgraded or long-lived sessions. A web server’s graceful close may cover ordinary requests while WebSockets, streams, or external load balancers need their own protocol.

03 / Define the contract

Choose what happens to new work and active work.

Two populations, two explicit behaviors
Request stateShutdown behaviorEvidence to collect
Not accepted yetMark the instance unready and stop admitting work when shutdown begins. During routing convergence, return a deliberate retryable response or close safely.Request arrival time, Pod UID, readiness and endpoint condition, response status.
Accepted and in flightAllow bounded work to finish. At the deadline, cancel or close according to the application contract.Request duration, completion/cancellation reason, drain deadline, client outcome.
Long-lived or upgradedDefine a separate reconnect, close-frame, or migration policy; ordinary HTTP server close behavior may not cover it.Connection type, session age, close code, reconnect and duplicate-action behavior.

Readiness indicates whether an instance should receive new work under the service’s routing contract. It is not a command that instantly empties every connection. A shutdown flag should be idempotent: Kubernetes lifecycle hooks can be retried, and repeated drain requests must not reopen service or extend the deadline without limit.

04 / Set the budget

Give the application room to finish, with a hard stop.

The example allows five seconds for the hook and up to twenty-five seconds for application shutdown inside a forty-second Pod grace period. Ten seconds remain for scheduling, signal delivery, and other termination overhead. Those are teaching values, not universal defaults: derive them from measured request durations, client deadlines, and the reliability cost of waiting.

If requests can validly take longer than the drain window, either revise the request contract or accept that some work will be cut off. Infinite waiting can strand capacity and block rollout progress. A deadline is a deliberate boundary, and its expiration should be observable.

deployment.yaml · termination budget
apiVersion: apps/v1
kind: Deployment
metadata:
  name: reports-api
spec:
  replicas: 3
  selector:
    matchLabels:
      app: reports-api
  template:
    metadata:
      labels:
        app: reports-api
    spec:
      # 5s preStop allowance + up to 25s active-request drain + 10s margin.
      terminationGracePeriodSeconds: 40
      containers:
        - name: api
          image: reports-api:reviewed-release
          ports:
            - name: http
              containerPort: 8080
          lifecycle:
            preStop:
              httpGet:
                path: /drain
                port: http
          readinessProbe:
            httpGet:
              path: /readyz
              port: http
            periodSeconds: 3
            timeoutSeconds: 1
            failureThreshold: 1
05 / Implement the drain

Make shutdown explicit in the server and bounded by a deadline.

Both examples set a drain state used by readiness, then stop accepting new connections and wait for ordinary in-flight work. The TypeScript version forces remaining connections closed after its deadline. The Go version passes a deadline context to http.Server.Shutdown and calls Close if graceful shutdown expires. Neither example automatically solves WebSocket or external load-balancer draining.

server.ts · stop, then drain
import { createServer } from 'node:http';
import { setTimeout as delay } from 'node:timers/promises';

let draining = false;

const server = createServer(async (request, response) => {
	if (request.url === '/readyz') {
		response.writeHead(draining ? 503 : 200).end();
		return;
	}
	if (request.url === '/drain') {
		draining = true;
		await delay(5_000); // bounded routing-propagation allowance for this example
		response.writeHead(200).end('draining');
		return;
	}
	if (draining) {
		response.writeHead(503, { connection: 'close' }).end('instance is draining');
		return;
	}
	await delay(3_000); // representative in-flight request
	response.writeHead(200).end('request complete');
});

server.listen(8080);

process.once('SIGTERM', () => {
	const forceStop = setTimeout(() => server.closeAllConnections(), 25_000);
	forceStop.unref();
	server.close(() => {
		clearTimeout(forceStop);
		process.exitCode = 0;
	});
});
server.go · bounded graceful shutdown
package main

import (
	"context"
	"log"
	"net/http"
	"os"
	"os/signal"
	"sync/atomic"
	"syscall"
	"time"
)

func main() {
	var draining atomic.Bool
	mux := http.NewServeMux()
	mux.HandleFunc("GET /readyz", func(w http.ResponseWriter, r *http.Request) {
		if draining.Load() {
			http.Error(w, "draining", http.StatusServiceUnavailable)
			return
		}
		w.WriteHeader(http.StatusOK)
	})
	mux.HandleFunc("GET /drain", func(w http.ResponseWriter, r *http.Request) {
		draining.Store(true)
		time.Sleep(5 * time.Second) // bounded routing-propagation allowance for this example
		w.WriteHeader(http.StatusOK)
	})
	mux.HandleFunc("GET /work", func(w http.ResponseWriter, r *http.Request) {
		if draining.Load() {
			http.Error(w, "instance is draining", http.StatusServiceUnavailable)
			return
		}
		select {
		case <-time.After(3 * time.Second):
			w.Write([]byte("request complete"))
		case <-r.Context().Done():
			return
		}
	})

	server := &http.Server{Addr: ":8080", Handler: mux}
	serveErr := make(chan error, 1)
	go func() { serveErr <- server.ListenAndServe() }()

	signals := make(chan os.Signal, 1)
	signal.Notify(signals, syscall.SIGTERM, os.Interrupt)
	select {
	case <-signals:
	case err := <-serveErr:
		if err != nil && err != http.ErrServerClosed {
			log.Fatal(err)
		}
		return
	}
	draining.Store(true)

	ctx, cancel := context.WithTimeout(context.Background(), 25*time.Second)
	defer cancel()
	if err := server.Shutdown(ctx); err != nil {
		log.Printf("grace period expired; close remaining connections: %v", err)
		_ = server.Close()
	}
	if err := <-serveErr; err != nil && err != http.ErrServerClosed {
		log.Fatal(err)
	}
}
06 / Verify the rollout

Exercise the transition and follow one request end to end.

In a disposable environment, run a request that lasts longer than a few seconds while deleting its Pod, then repeat with a request near the configured deadline. Record whether it completed, was cancelled, or was retried. Also issue fresh requests during the transition. Repeat with a persistent connection if the service uses them.

Compare application timestamps to Pod events and EndpointSlice state, then verify the client-visible result. A port-forward-only test skips much of the Service routing path; it can validate the process behavior but not the whole traffic transition. Observe failed and successful requests so a temporary error spike is not hidden by the final rollout status.

Change one variable per drill: try a shorter hook wait, a longer in-flight request, then a deadline expiry. Which observation would distinguish routing convergence from a process that closed active work?
07 / Keep the operating rule

Graceful means predictable completion or predictable cancellation.

Before shipping a shutdown change, name the request classes the service promises to finish, the deadline for each, the client retry and idempotency behavior, and the evidence that demonstrates the result. If the measured failures persist, keep the cause open: routing propagation, application close behavior, load-balancer policy, and client retries remain distinct places to investigate.

Try a decision

Follow the remaining resets

A rollout still shows occasional resets after the drain handler is added. What is the most useful next move?

Your next move

Operational rule to carry forward: stop admitting new work, finish accepted work only within a measured budget, and verify what users and clients actually observe across the complete request path.

References: Kubernetes container lifecycle hooks, Pod lifecycle, Node.js HTTP server close, and Go HTTP server shutdown. Accessed 2026-10-01.