Production debugging

Backend error and symptom index

Every page here covers one real failure: what the runtime is actually telling you, the causes ranked by how often they are the cause, the commands that prove which one you have, and the fix. Where Gronex ships a repository challenge that reproduces the failure, the page links to it — because reading about a production bug and debugging one are different skills.

Java / JVM

HikariPool-1 - Connection is not available, request timed out

Why HikariCP times out waiting for a connection, how to tell a leak from genuine saturation using pg_stat_activity and leakDetectionThreshold, and how to fix both.

java.lang.OutOfMemoryError: Java heap space

Distinguish a genuine memory leak from an unbounded single request, read a heap dump with jmap and MAT, and fix Java heap space OutOfMemoryError properly instead of raising -Xmx.

java.lang.OutOfMemoryError: Metaspace

Metaspace OutOfMemoryError almost always means classes are being loaded and never unloaded. How to count loaded classes, find the classloader leak, and fix it.

java.lang.OutOfMemoryError: GC overhead limit exceeded

GC overhead limit exceeded means the JVM spent over 98% of its time collecting and recovered under 2% of the heap. How to read the GC log and fix the cause, not the threshold.

java.util.ConcurrentModificationException

ConcurrentModificationException is usually single-threaded: structural modification during iteration. How fail-fast iterators detect it, the four correct fixes, and why synchronizing does not help.

OutOfMemoryError: unable to create new native thread

This is not a heap problem. The OS refused a new thread because of a thread leak, an unbounded executor, or a low pid/nproc limit. How to count threads and find the leak.

RejectedExecutionException: Task rejected from ThreadPoolExecutor

Read the executor state in the message, tell saturation from a shutdown pool, and choose the right rejection policy instead of making the queue unbounded.

PostgreSQL

FATAL: sorry, too many clients already

PostgreSQL refused a connection because max_connections is reached. How to find which service is hoarding connections, why bigger pools make it worse, and when to add a pooler.

Sessions stuck in "idle in transaction"

Why idle in transaction sessions block vacuum, hold locks, and exhaust your pool — how to find the code path that leaves them open, and the timeouts that contain the damage.

ERROR: deadlock detected

PostgreSQL rolled back one transaction to break a lock cycle. How to read the DETAIL, why unordered multi-row updates cause it, and the fixes that actually remove the cycle.

ERROR: canceling statement due to statement timeout

The timeout is the messenger. How to tell a missing index from lock waiting from an unbounded result set, using EXPLAIN ANALYZE and pg_stat_statements.

ERROR: could not serialize access due to concurrent update

PostgreSQL aborted your transaction to preserve isolation. Why REPEATABLE READ and SERIALIZABLE produce this, when a retry is the correct fix, and when to lock instead.

ERROR: current transaction is aborted, commands ignored until end of transaction block

This is a follow-on error, not the real one. How to find the original failure it is hiding, why savepoints matter, and how ORMs produce it.

ERROR: duplicate key value violates unique constraint

Sometimes this error is the constraint saving you from a double charge. How to tell a retry collision from a real bug, and why ON CONFLICT beats check-then-insert.

Deep OFFSET pagination getting slower every page

OFFSET makes PostgreSQL produce and discard every skipped row, so the last page is the slowest. How keyset pagination makes page cost constant, and its trade-offs.

ERROR: canceling statement due to lock timeout (ALTER TABLE)

Why a one-second ALTER TABLE can take your site down, how the ACCESS EXCLUSIVE lock queue blocks every reader behind it, and the lock_timeout pattern that makes migrations safe.

ERROR: canceling statement due to conflict with recovery

A standby cancelled your query because replay needed to remove rows it was reading. The trade-off between replica lag and query cancellation, and how to choose.

Concurrency

Found one Java-level deadlock (thread dump)

The JVM already found the cycle for you. How to read a deadlock in a thread dump, why inconsistent lock ordering causes it, and the four fixes ranked by robustness.

Threads blocked forever on a ReentrantLock with no deadlock reported

Why an exception between lock() and try leaves a ReentrantLock held forever, how to spot a leaked lock in a thread dump when the JVM reports no deadlock, and the correct acquire shape.

ThreadLocal value leaking across requests on a pooled thread

Why a ThreadLocal set on a pool thread outlives the request, how the weak key and strong value actually work, and how the same bug produces both data bleed and a heap leak.

Consumer stuck in Object.wait() with work already in the queue

Why a notify() that arrives before wait() is lost forever, why spurious wakeups are legal, and why the condition must be re-checked in a loop rather than tested once.

java.lang.IllegalMonitorStateException: current thread is not owner

Why wait(), notify() and unlock() demand ownership, what "current thread is not owner" actually proves about your code, and the two structural bugs that produce it.

Thread pool starvation — every worker waiting on a task in its own pool

A pooled task that submits to its own pool and waits deadlocks the pool without any lock. How to recognise it in a thread dump and restructure so it cannot happen.

Partially constructed object published by double-checked locking

The exact reordering that makes non-volatile double-checked locking broken, why it passes thousands of runs before failing, and the three correct lazy-initialisation idioms.

Lost update from get-then-put on a ConcurrentHashMap

Why a thread-safe map does not make your read-modify-write thread-safe, why compute() and merge() are atomic while get-then-put is not, and how to pick the right atomic operation.

CompletableFuture failed with nothing logged

Why a failed CompletableFuture stage produces no log line, how exceptionally, handle and whenComplete differ, and how to make async failures impossible to lose.

awaitTermination never returns and the JVM will not exit

Why shutdown() waits for running tasks, why shutdownNow() cannot stop a socket read, and the shutdown sequence that terminates reliably inside a SIGTERM grace period.

InterruptedException caught and ignored — the task can no longer be cancelled

Catching InterruptedException clears the interrupt flag. Why that silently disables cancellation, timeouts and shutdown, and the two correct ways to handle it.

Two unrelated components sharing a monitor via a boxed Integer or interned String

Why locking on Integer.valueOf, a String literal, or Boolean shares a monitor JVM-wide, how the autobox cache and string pool cause it, and the safe alternatives.

Cache stampede — the same expensive value built many times concurrently

Why a cache miss under concurrency triggers a thundering herd of identical loads, why synchronising the whole cache is the wrong fix, and the future-per-key memoiser that works.

Lock convoy — throughput collapses as threads are added, with no deadlock

Why a coarse lock on a hot path collapses throughput without deadlocking, how to tell a convoy from a deadlock and from livelock, and how to fix it structurally.

ReadWriteLock writer blocked indefinitely behind a stream of readers

Why a non-fair ReentrantReadWriteLock can block a writer indefinitely under continuous reads, why read-to-write upgrade self-deadlocks, and the alternatives that do not starve.

Worker loop never sees the stop flag and runs forever

Why a worker thread can loop forever on a boolean another thread already set, what the JMM actually guarantees, and why volatile is the fix and sleep() is not.

Distributed systems

CommitFailedException: Commit cannot be completed since the group has already rebalanced

Why Kafka revokes your partitions mid-batch, how max.poll.interval.ms and max.poll.records interact, and how to stop the duplicate processing this exception always leaves behind.

Consumer group stuck rebalancing — poll timeout has expired

A consumer group that rebalances forever makes no progress. How generation churn, range assignment, and one slow member create the loop, and how to break it.

The same message processed twice (at-least-once delivery)

Every mainstream broker delivers at least once. Why redelivery is guaranteed rather than exceptional, and the storage-level patterns that make reprocessing harmless.

Messages processed out of order across partitions

Ordering in Kafka is per partition, never per topic. How a key change or a partition-count change reorders events, and the version-guard pattern that makes order irrelevant.

Webhook delivered twice — customer charged twice

Why payment providers retry webhooks, why a slow handler guarantees duplicates, and how to build a webhook endpoint whose effects are safe under redelivery and reordering.

Database and broker diverge after a dual write

Why writing to a database and a broker in one method can never be atomic, why the compensating-catch fix fails, and how the transactional outbox pattern actually closes the gap.

Retry storm: thundering herd after a dependency failure

Diagnose synchronized retries that amplify an outage, then apply bounded exponential backoff, jitter, and retry budgets.

Outbound call has no timeout and exhausts workers

Trace a cascading failure from one hung dependency through worker and pool exhaustion, then set deadlines that compose.

Distributed lock lease expired while the holder was still working

Why a pause lets an old lock holder keep writing, and how lease renewal, fencing tokens, and ownership checks prevent corruption.

Clock skew: timestamp ordering or token expiry is inconsistent

Find wall-clock assumptions that fail across hosts and replace them with monotonic time, server timestamps, or logical ordering.

Exactly-once claim fails at an external side effect

Separate broker exactly-once transactions from external side effects, then build effectively-once outcomes with idempotency.

Dead-letter queue growing without an alert

Treat a growing DLQ as a correctness incident: identify poison messages, preserve ordering, and replay after fixing the cause.

Idempotency key reused with a different request body

Explain why an idempotency key must bind to request intent, how to fingerprint canonical bodies, and how to handle legitimate retries.

Read-after-write returned stale data from a replica

Diagnose stale reads after a successful write and use session guarantees, LSN waiting, or primary routing deliberately.

Replication lag: replica served stale data after a write

Measure write-ahead-log lag, find the slow replay stage, and choose routing or backpressure before stale data becomes corruption.

502 Bad Gateway from a reverse proxy

Diagnose reverse-proxy 502 responses using upstream timing, direct-origin probes and nginx error logs. Separate connection refusal, invalid headers and crashes.

504 Gateway Timeout

Find which proxy timer produced a 504, distinguish connection and read waits, and propagate deadlines so timed-out requests stop consuming backend capacity.

CORS preflight: missing Access-Control-Allow-Origin

Diagnose missing Access-Control-Allow-Origin on preflight, authentication failures and cached responses. Fix exact origins, credentials and error-path headers.

413 Payload Too Large

Trace a 413 through CDN, nginx and application parsers. Measure encoded body size, scope upload limits and keep memory bounded for large or compressed requests.

429 Too Many Requests and Retry-After

Handle HTTP 429 with delay-seconds or HTTP-date Retry-After, shared quota coordination and a bounded retry budget. Avoid amplifying a rate-limit incident.

nginx 499 client closed request

Explain nginx 499 access-log entries, identify client or outer-proxy deadlines, and stop abandoned requests from holding application and database capacity.

upstream prematurely closed connection

Trace nginx premature upstream closure to crashes, worker timeouts, deployment draining or protocol mistakes. Distinguish header failures from truncated bodies.

Intermittent 502 after an idle keep-alive connection

Diagnose stale upstream connection reuse using idle-gap reproduction and connection timing. Align proxy and origin keep-alive lifetimes without unsafe write retries.

SSL certificate problem: unable to get local issuer certificate

Use curl and OpenSSL to separate a missing intermediate certificate from an absent trust root, enterprise TLS interception and the wrong client CA bundle.

ERR_INCOMPLETE_CHUNKED_ENCODING

Diagnose truncated HTTP/1.1 chunked responses despite status 200. Compare direct-origin streams, nginx logs and client transfer completion before retrying.

Request timeouts cascading into pool exhaustion

Explain why timed-out HTTP requests still occupy database or connection pools. Diagnose queueing, cancellation and retries, then bound each resource lifetime.

Connection reset by peer on a long-polling endpoint

Find which hop resets long-poll requests, distinguish active-request idle limits from keep-alive, and design bounded polls with cursors and cancellation.

Redis OOM command not allowed above maxmemory

Separate Redis write rejection from a host OOM kill. Inspect memory policy, TTL eligibility and oversized keys before deciding whether eviction is safe for your data.

MISCONF Redis is configured to save RDB snapshots

Find why Redis background snapshots failed using INFO persistence and server logs. Distinguish storage, permissions and fork pressure without discarding durability policy.

READONLY You can’t write against a read only replica

Identify the Redis node actually receiving writes, distinguish reader endpoints from stale failover connections, and restore primary discovery without enabling replica writes.

LOADING Redis is loading the dataset in memory

Determine whether Redis loading is progressing or restarting. Inspect persistence progress, uptime and storage before tuning readiness checks and bounded client retries.

CROSSSLOT Keys in request don’t hash to the same slot

Prove slot mismatches with CLUSTER KEYSLOT, choose hash tags around real transaction boundaries, and avoid turning a CROSSSLOT fix into a hot shard or lost atomicity.

MOVED and ASK replies from Redis Cluster

Handle Redis Cluster redirects correctly: refresh MOVED slot ownership, pair ASKING with the redirected request on one connection, and validate advertised node addresses.

Redis clients stall during KEYS on a large keyspace

Find blocking keyspace scans with Redis SLOWLOG and command statistics. Replace request-path enumeration with bounded SCAN or an explicit index without assuming snapshot semantics.

A Redis lock lease expires while the worker still runs

Reconstruct overlapping Redis lease holders, prevent stale unlocks with owner tokens, and enforce correctness at the protected resource when a paused worker resumes.

Python

psycopg2.InterfaceError: connection already closed

The connection object is alive in Python but the socket is gone. How to tell a server-side termination from your own double close, and how pre-ping and recycle actually fix it.

RecursionError: maximum recursion depth exceeded

Find the missing base case, accidental object cycle, or pathological input behind Python RecursionError and fix the algorithm rather than merely raising the limit.

MemoryError: unable to allocate array

Diagnose Python MemoryError by measuring result cardinality and resident memory, then stream, project, chunk, and aggregate data safely.

QueuePool limit of size 5 overflow 10 reached, connection timed out

QueuePool timeouts mean sessions are not being returned. How to tell a leaked session from genuine saturation, using pool events and pg_stat_activity, and fix the lifecycle.

DetachedInstanceError: instance is not bound to a Session

DetachedInstanceError means an ORM object outlived its Session and then needed the database. Why expire_on_commit causes it, and the four correct fixes ranked.

RuntimeError: Event loop is closed

Event loop is closed usually means an object outlived the loop that created it, or asyncio.run was called more than once. How to find the owner and fix the lifecycle.

Task was destroyed but it is pending!

This warning means asyncio work vanished without completing — a garbage-collected task or an abrupt shutdown. Why fire-and-forget tasks lose data, and how to structure them safely.

Executing <Handle ...> took 2.418 seconds (blocked event loop)

One synchronous call inside async def stalls every concurrent request. How to measure event-loop lag, find the blocking frame, and move the work off the loop correctly.

SettingWithCopyWarning: A value is trying to be set on a copy of a slice

Understand pandas view versus copy ambiguity, prove whether an assignment reached the source frame, and make chained indexing deterministic.

celery.exceptions.WorkerLostError: Worker exited prematurely

Separate Celery worker crashes, SIGKILL from the OOM killer, and hard time limits using worker and kernel evidence.

requests.exceptions.ReadTimeout: HTTPSConnectionPool read timed out

Distinguish DNS/connect/TLS delay from slow response bytes in requests, then set phase-aware timeouts and bounded retries.

UnicodeDecodeError: 'utf-8' codec can't decode byte

Diagnose byte-versus-text boundaries and identify encodings safely instead of hiding corrupted input with errors=ignore.

AssertionError: daemonic processes are not allowed to have children

Why a multiprocessing pool worker cannot create a child process, and how to choose threads, a non-daemonic supervisor, or a flat process topology.

[CRITICAL] WORKER TIMEOUT (pid:1234)

Find whether a Gunicorn worker is blocked on CPU, I/O, or a downstream call, then align timeouts and worker models without masking the cause.

Mutable default argument retains state across calls

Explain why Python evaluates defaults once, identify cross-request state leaks, and repair mutable defaults without masking concurrency issues.

Node.js

FATAL ERROR: Reached heap limit Allocation failed

Diagnose Node.js heap exhaustion with memory samples and heap snapshots. Separate retained objects, oversized requests and external buffers before changing heap limits.

Unhandled promise rejection crashes the process

Find promises detached from request error handling, understand Node rejection policy and repair async callbacks, finally chains and background task ownership.

ECONNRESET: socket hang up on a reused connection

Identify ECONNRESET on reused Node HTTP sockets, distinguish idle-close races from upstream failure, and retry only when the operation is safe to repeat.

MaxListenersExceededWarning: possible EventEmitter memory leak

Trace listener registration growth, distinguish intentional fan-out from request leaks, and clean up EventEmitter subscriptions on success, failure and disconnect.

EADDRINUSE: address already in use

Diagnose Node EADDRINUSE using listening sockets and process ancestry. Fix duplicate startup, container namespace confusion and test-server teardown safely.

ERR_HTTP_HEADERS_SENT

Find the first response commit behind ERR_HTTP_HEADERS_SENT. Repair missing returns, competing async branches and stream error handling without hiding the race.

Event loop blocked by synchronous work

Measure event-loop delay in milliseconds, correlate CPU profiles and request latency, and distinguish synchronous JavaScript, blocking I/O and worker-pool contention.

pg client already connected or released twice

Distinguish pg Client and Pool lifecycles, avoid released-client reuse, and keep transaction queries on one checked-out connection with unconditional cleanup.

Sequelize / Knex pool acquire timeout

Find why Sequelize or Knex cannot acquire a connection: transaction context escapes, leaked borrowers, slow SQL and excess concurrency. Fix ownership before pool size.

ERR_STREAM_PREMATURE_CLOSE during an upload

Diagnose a piped upload closing before completion. Find the first stream failure, distinguish client aborts from storage errors and publish files only after validation.

Process exits before asynchronous writes finish

Understand why process.exit truncates output, why promises alone do not keep Node alive, and how to drain servers, streams and database pools before shutdown.

ERR_MODULE_NOT_FOUND during ESM/CommonJS migration

Diagnose missing ESM imports using the emitted file, explicit extensions, package exports and production dependencies. Separate resolution failures from interop errors.

Unbounded JSON body blocks the loop or exhausts memory

Prevent oversized JSON requests from exhausting Node memory or blocking the event loop. Enforce byte limits before buffering, account for decompression and bound concurrency.

Undici / fetch connections remain occupied

Diagnose fetch queues and socket pressure caused by unread response bodies, per-request dispatchers and unlimited concurrency. Fix body ownership and dispatcher lifecycle.

ERR_UNHANDLED_REJECTION in a worker thread

Repair unhandled async worker callbacks, distinguish task failures from worker crashes, and settle parent promises once when a worker errors or exits.

Linux / shell

command not found in a script that works interactively

Trace why cron, systemd or CI cannot find a command that works in your terminal. Separate PATH lookup, shell functions, runtime activation and working-directory errors.

Permission denied when executing a script

Diagnose execution permission failures using namei, findmnt and the real service identity. Distinguish missing execute bits, inaccessible interpreters and noexec mounts.

bad interpreter: No such file or directory with CRLF

Expose the carriage return hidden in a Linux shebang, distinguish it from a missing interpreter or ELF loader, and repair line endings without changing script behaviour.

Argument list too long

Fix Linux argument-list limits without losing filenames containing spaces or newlines. Understand glob expansion, environment size and why xargs with a glob still fails.

Too many open files

Measure the failing process’s descriptor count and limits, separate leaks from bounded concurrency, and fix file or socket ownership before changing LimitNOFILE.

Out of memory: Killed process

Investigate a SIGKILL using kernel logs and cgroup v2 memory.events. Separate container limits, host pressure and application allocation failures before sizing memory.

No space left on device despite free disk space

Use df -i on the actual target filesystem to distinguish inode exhaustion from full blocks, wrong mounts and deleted open files. Fix small-file growth at its source.

Text file busy during executable replacement

Understand ETXTBSY when overwriting a running executable or executing a file still open for writing. Use immutable release artifacts and same-filesystem rename.

set -e script continues after a failed pipeline

Reproduce why Bash errexit misses upstream pipeline failures, conditional functions and command substitutions. Use pipefail plus explicit checks at publication boundaries.

An unquoted variable turns one argument into several

See the exact argument vector produced by Bash expansion. Preserve paths with quotes, use arrays for command arguments, and handle empty values and leading dashes explicitly.

OOMKilled — container exit code 137

Distinguish a container memory-limit kill from node pressure and other SIGKILL exits. Inspect termination state, cgroup counters and peak allocation before tuning limits.

CrashLoopBackOff

Trace CrashLoopBackOff to the previous container exit, probe failure or completed foreground process. Preserve evidence and repair the cause of repeated restarts.

ImagePullBackOff / ErrImagePull

Read the registry failure behind ImagePullBackOff. Separate missing images, pull-secret scope, node egress, rate limits and incompatible image platforms.

Readiness probe failed

Compare the kubelet probe with the request that succeeds. Check pod IP binding, port, HTTP contract, timeouts and dependency checks before changing readiness.

Liveness probe failed during startup

Prove that kubelet interrupts legitimate initialisation, separate startup from liveness, and size a startup probe using cold-start measurements under real limits.

CreateContainerConfigError — missing Secret or ConfigMap

Find the exact unresolved configuration reference before container startup. Check namespace, key names, generated resource names and secret-controller reconciliation.

Evicted — node disk pressure

Distinguish node DiskPressure from a pod storage-limit eviction. Trace logs, writable layers, emptyDir and inode consumption without deleting runtime state.

runAsNonRoot and image will run as root

Resolve Kubernetes runAsNonRoot validation with an explicit numeric UID and compatible file ownership. Distinguish UID validation from application permission failures.

JVM container killed despite available Java heap

Understand why MaxRAMPercentage and Xmx do not cap total Java process memory. Compare cgroup usage, effective JVM flags, thread stacks and native allocation.

exec format error — container architecture mismatch

Trace exec format error through node architecture, image manifests and the actual executable. Fix cross-builds and distinguish ELF mismatch from broken entrypoint scripts.

no space left on device — container layers

Diagnose ENOSPC in image pulls, builds and writable layers. Understand OverlayFS copy-up, immutable lower layers, inodes and Docker daemon storage limits.

SIGTERM misses the application — shutdown ends in SIGKILL

Trace shutdown signals through shell wrappers, PID 1 and application handlers. Use exec correctly, budget preStop time and test draining with work in flight.

Debug these for real, not from memory

Gronex problems are broken backend repositories with failing test suites that encode the production invariant — leaked connections, unbounded caches, lock cycles, lost messages. You get the evidence an on-call engineer would get, and the tests pass only when the invariant actually holds.