← Networking
Technique Operations and troubleshooting

Timeouts, keepalives, and connection pooling

A timeout only helps when you know which wait it bounds.

A service slows down in bursts. Some requests wait for an available pooled connection; some reach a peer that has already forgotten an idle flow; and a retry adds more work. Trace one request through its waits, then decide which deadline, pool limit, or connection policy applies.

Measure each wait

Keep the request inside one end-to-end budget, and distinguish connection reuse from transport probes.

TypeScriptGo Deadlines · connection reuse · pool limits
01 / Find where time went

A request can wait before it ever opens a connection.

An outbound request can wait for a pool slot, name resolution, TCP establishment, TLS negotiation, request upload, response headers, and response body. A metric called “request duration” may include all of these, but a connect timer or header timer covers only part of the path.

Start with a timeline from the caller’s perspective. Record when the request entered the client, acquired a connection, connected, sent its request, received headers, and finished the body. Then compare it with server and proxy timings. A server that reports a fast handler did not measure time spent in the caller’s pool or network path.

Before a socketPool queue

Wait for an allowed connection or free slot.

Open a pathDNS · TCP · TLS

Resolve and establish transport and security.

Ask the serviceWrite · headers

Send the request and wait for a response to begin.

Finish the resultRead body

Consume data until completion or the deadline.

02 / Give the work one deadline

Phase limits sit inside the caller’s remaining time.

Give the operation an overall deadline based on the caller’s budget. Use shorter phase limits to catch a stalled connect or a slow response, while preserving time for later work. Every nested dependency should receive only the time remaining from its caller—not a fresh full budget.

For example, a 700 ms caller budget could allow up to 150 ms for a new connection and 300 ms for response headers, leaving the rest for pool wait, transfer, and cleanup. These are illustrative numbers, not safe defaults. Set them from latency objectives, measurements, and the cost of abandoning work.

Illustrative request budget · 700 ms totalPhase limits cannot extend the parent deadline
Queue + setup
At most 150 ms for a new TCP connection; pool wait still consumes the total budget.
Response start
Up to 300 ms for headers after the request is written.
Transfer + cleanup
Use the remaining time; cancel downstream work when the caller’s deadline expires.
03 / Account for pool queues

A connection limit is also a queueing limit.

A connection pool reuses established connections and caps how many it opens. If all slots are busy, another request waits. That wait can consume its deadline even though it has not reached DNS, TCP, or the server. A small pool protects the peer and host from connection churn; a pool that is too small for the workload can create queueing delay.

Observe active, idle, and queued requests alongside request latency. Compare pool wait with service time and connection reuse. Increasing the maximum may reduce a local queue but also raise downstream concurrency, file descriptor use, and load. Change one constraint at a time and watch the peer’s capacity.

Use evidence at the layer where the wait occurs
SignalMay indicateCheck next
High pool wait, low server timeConnection slots are occupied or pool capacity is constrained.Active requests, configured limits, and whether responses are fully consumed.
Long connect phaseResolution or transport establishment is delayed or filtered.Resolve time separately; inspect route, handshake, and connection errors.
Headers are prompt, body stallsStreaming response paused or peer stopped producing data.Read progress, body size, inactivity policy, and total remaining budget.
04 / Name the kind of keepalive

Reuse, TCP probes, and application heartbeats do different jobs.

HTTP connection persistence lets a client reuse a TCP connection for another request. TCP keepalive is an optional transport probe that can check an otherwise idle connection after operating-system-configured intervals. An application heartbeat is a protocol message that asks whether the remote application is responsive.

None guarantees that every proxy, NAT, firewall, or peer will retain the flow. An intermediary may expire idle state before the client probes; a heartbeat can prove an application response only if the protocol defines one; and a TCP acknowledgment does not prove that the remote application is healthy. Set idle policies with the actual path and runtime in mind.

05 / Observe reuse and timeout

Compare peer ports with one pooled local connection.

These examples run an HTTP server and client in the same process, bound to 127.0.0.1. The pool allows one connection: the second request waits while the first is slow, then reuses that connection. The final delayed response triggers a short socket-inactivity limit in TypeScript and a response-header timeout in Go. No remote system is contacted.

Read each API in context. The TypeScript example sets an overall timer, a connection timer for a newly created socket, and a socket inactivity timer. The Go transport limits connections per host and exposes dial and response-header limits; the context and client timeouts bound a larger portion of the request. Pool wait and deadline scope should be verified against the library documentation.

Compare the same pool scenario in TypeScript and Go.

Both clients cap the local pool at one connection and put bounds on slow work.

TypeScriptLocal HTTP pool · queue, reuse, and a bounded slow response
pool.ts
import { Agent, createServer, request } from 'node:http';
import type { AddressInfo } from 'node:net';

type Limits = { connectMs: number; idleMs: number; totalMs: number };

function getText(port: number, agent: Agent, path: string, limits: Limits): Promise<string> {
	return new Promise((resolve, reject) => {
		const req = request({ host: '127.0.0.1', port, path, agent }, (response) => {
			response.setEncoding('utf8');
			let body = '';
			response.on('data', (chunk: string) => (body += chunk));
			response.on('end', () => {
				clearTimeout(totalTimer);
				if (connectTimer) clearTimeout(connectTimer);
				resolve(body.trim());
			});
		});

		const totalTimer = setTimeout(
			() => req.destroy(new Error('total request deadline exceeded')),
			limits.totalMs
		);
		let connectTimer: ReturnType<typeof setTimeout> | undefined;
		req.setTimeout(limits.idleMs, () => req.destroy(new Error('socket inactivity limit exceeded')));
		req.on('socket', (socket) => {
			if (!socket.connecting) return; // A reused connection has no new connect phase.
			connectTimer = setTimeout(
				() => socket.destroy(new Error('TCP connect deadline exceeded')),
				limits.connectMs
			);
			socket.once('connect', () => {
				if (connectTimer) clearTimeout(connectTimer);
			});
			socket.once('close', () => {
				if (connectTimer) clearTimeout(connectTimer);
			});
		});
		req.once('error', (error) => {
			clearTimeout(totalTimer);
			if (connectTimer) clearTimeout(connectTimer);
			reject(error);
		});
		req.end();
	});
}

async function main() {
	// The only origin is this process, bound to IPv4 loopback.
	const server = createServer((req, response) => {
		const peerPort = req.socket.remotePort;
		const delay = req.url === '/slow' ? 140 : req.url === '/very-slow' ? 250 : 0;
		setTimeout(() => {
			if (response.destroyed) return;
			response.writeHead(200, { 'content-type': 'text/plain' });
			response.end(`path=${req.url} peer-port=${peerPort}\n`);
		}, delay);
	});
	server.keepAliveTimeout = 1000;
	await new Promise<void>((resolve, reject) => {
		server.once('error', reject);
		server.listen(0, '127.0.0.1', resolve);
	});

	const address = server.address();
	if (!address || typeof address === 'string') throw new Error('Expected a local TCP address.');
	const agent = new Agent({ keepAlive: true, maxSockets: 1, maxFreeSockets: 1 });
	const port = (address as AddressInfo).port;
	const normalLimits = { connectMs: 500, idleMs: 500, totalMs: 1000 };

	try {
		// The second request waits in the Agent queue while /slow owns the one socket.
		const [slow, fast] = await Promise.all([
			getText(port, agent, '/slow', normalLimits),
			getText(port, agent, '/fast', normalLimits)
		]);
		console.log(slow);
		console.log(fast);

		try {
			await getText(port, agent, '/very-slow', {
				connectMs: 500,
				idleMs: 60,
				totalMs: 500
			});
		} catch (error) {
			console.log(`bounded slow request: ${(error as Error).message}`);
		}
	} finally {
		agent.destroy();
		await new Promise<void>((resolve) => server.close(() => resolve()));
	}
}

void main().catch((error: unknown) => {
	console.error(error);
	process.exitCode = 1;
});
GoLocal HTTP pool · queue, reuse, and a bounded slow response
pool.go
package main

import (
	"context"
	"fmt"
	"io"
	"net"
	"net/http"
	"time"
)

func main() {
	started := make(chan struct{}, 1)
	mux := http.NewServeMux()
	mux.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
		delay := time.Duration(0)
		switch r.URL.Path {
		case "/slow":
			delay = 140 * time.Millisecond
			select {
			case started <- struct{}{}:
			default:
			}
		case "/very-slow":
			delay = 500 * time.Millisecond
		}
		time.Sleep(delay)
		_, _ = fmt.Fprintf(w, "path=%s peer=%s\n", r.URL.Path, r.RemoteAddr)
	})

	listener, err := net.Listen("tcp", "127.0.0.1:0")
	if err != nil {
		panic(err)
	}
	server := &http.Server{Handler: mux}
	go func() { _ = server.Serve(listener) }()
	defer server.Close()

	transport := &http.Transport{
		MaxConnsPerHost:       1,
		MaxIdleConnsPerHost:   1,
		IdleConnTimeout:       300 * time.Millisecond,
		ResponseHeaderTimeout: 200 * time.Millisecond,
		DialContext: (&net.Dialer{
			Timeout: 100 * time.Millisecond,
		}).DialContext,
	}
	client := &http.Client{Transport: transport, Timeout: time.Second}
	defer transport.CloseIdleConnections()

	get := func(path string, budget time.Duration) (string, error) {
		ctx, cancel := context.WithTimeout(context.Background(), budget)
		defer cancel()
		req, err := http.NewRequestWithContext(ctx, http.MethodGet, "http://"+listener.Addr().String()+path, nil)
		if err != nil {
			return "", err
		}
		response, err := client.Do(req)
		if err != nil {
			return "", err
		}
		defer response.Body.Close()
		body, err := io.ReadAll(response.Body)
		return string(body), err
	}

	type result struct {
		body string
		err  error
	}
	slowResult := make(chan result, 1)
	go func() {
		body, err := get("/slow", 900*time.Millisecond)
		slowResult <- result{body: body, err: err}
	}()
	<-started // Make the slower request occupy the only connection before queuing /fast.
	fast, err := get("/fast", 900*time.Millisecond)
	if err != nil {
		panic(err)
	}
	slow := <-slowResult
	if slow.err != nil {
		panic(slow.err)
	}
	fmt.Print(slow.body)
	fmt.Print(fast)

	_, err = get("/very-slow", 900*time.Millisecond)
	fmt.Printf("bounded slow request: %v\n", err)
}
06 / Choose a safe repair

Bound concurrency, cancel expired work, and retry with a budget.

Use a caller deadline, propagate cancellation to downstream work, and set phase limits that fit inside the remaining budget. Size pools against measured concurrency and the downstream service’s capacity. Make sure response bodies are consumed or closed so connections can be returned to the pool. Align idle reuse with known path policies where those are documented.

A timeout does not tell you whether a write or side effect reached the server. Before retrying, establish whether the operation is idempotent or protected by an idempotency key, and cap attempts by the remaining deadline. Backoff and jitter can reduce synchronized retry pressure, but retrying an overloaded dependency without a budget can amplify the incident.

07 / Carry the model forward

Choose metrics and settings from the actual client contract.

RFC 1122, Requirements for Internet Hosts describes TCP keepalive behavior and leaves probes optional. The client APIs document their specific timer and pool behavior: Node.js node:http and Go net/http. Verify details for the deployed runtime version; defaults and paths vary.

For the retry side of an incident, use the linked Math in Practice lesson on retry amplification and backoff. A deadline limits caller work; it does not by itself make a retry safe.