Gronex
Log in

Production Errors

120 guides

What a runtime error actually means, how to prove the cause, and how to fix it.

HikariPool-1 - Connection is not available, request timed outWhy HikariCP times out waiting for a connection, how to tell a leak from genuine saturation using pg_stat_activity and leakDetectionThreshold, and how to fix both.Java / JVMjava.lang.OutOfMemoryError: Java heap spaceDistinguish a genuine memory leak from an unbounded single request, read a heap dump with jmap and MAT, and fix Java heap space OutOfMemoryError properly instead of raising -Xmx.Java / JVMjava.lang.OutOfMemoryError: MetaspaceMetaspace OutOfMemoryError almost always means classes are being loaded and never unloaded. How to count loaded classes, find the classloader leak, and fix it.Java / JVMjava.lang.OutOfMemoryError: GC overhead limit exceededGC overhead limit exceeded means the JVM spent over 98% of its time collecting and recovered under 2% of the heap. How to read the GC log and fix the cause, not the threshold.Java / JVMjava.util.ConcurrentModificationExceptionConcurrentModificationException is usually single-threaded: structural modification during iteration. How fail-fast iterators detect it, the four correct fixes, and why synchronizing does not help.Java / JVMOutOfMemoryError: unable to create new native threadThis is not a heap problem. The OS refused a new thread because of a thread leak, an unbounded executor, or a low pid/nproc limit. How to count threads and find the leak.Java / JVMRejectedExecutionException: Task rejected from ThreadPoolExecutorRead the executor state in the message, tell saturation from a shutdown pool, and choose the right rejection policy instead of making the queue unbounded.Java / JVMFATAL: sorry, too many clients alreadyPostgreSQL refused a connection because max_connections is reached. How to find which service is hoarding connections, why bigger pools make it worse, and when to add a pooler.PostgreSQLSessions stuck in "idle in transaction"Why idle in transaction sessions block vacuum, hold locks, and exhaust your pool — how to find the code path that leaves them open, and the timeouts that contain the damage.PostgreSQLERROR: deadlock detectedPostgreSQL rolled back one transaction to break a lock cycle. How to read the DETAIL, why unordered multi-row updates cause it, and the fixes that actually remove the cycle.PostgreSQLERROR: canceling statement due to statement timeoutThe timeout is the messenger. How to tell a missing index from lock waiting from an unbounded result set, using EXPLAIN ANALYZE and pg_stat_statements.PostgreSQLERROR: could not serialize access due to concurrent updatePostgreSQL aborted your transaction to preserve isolation. Why REPEATABLE READ and SERIALIZABLE produce this, when a retry is the correct fix, and when to lock instead.PostgreSQLERROR: current transaction is aborted, commands ignored until end of transaction blockThis is a follow-on error, not the real one. How to find the original failure it is hiding, why savepoints matter, and how ORMs produce it.PostgreSQLERROR: duplicate key value violates unique constraintSometimes this error is the constraint saving you from a double charge. How to tell a retry collision from a real bug, and why ON CONFLICT beats check-then-insert.PostgreSQLDeep OFFSET pagination getting slower every pageOFFSET makes PostgreSQL produce and discard every skipped row, so the last page is the slowest. How keyset pagination makes page cost constant, and its trade-offs.PostgreSQLERROR: canceling statement due to lock timeout (ALTER TABLE)Why a one-second ALTER TABLE can take your site down, how the ACCESS EXCLUSIVE lock queue blocks every reader behind it, and the lock_timeout pattern that makes migrations safe.PostgreSQLERROR: canceling statement due to conflict with recoveryA standby cancelled your query because replay needed to remove rows it was reading. The trade-off between replica lag and query cancellation, and how to choose.PostgreSQLFound one Java-level deadlock (thread dump)The JVM already found the cycle for you. How to read a deadlock in a thread dump, why inconsistent lock ordering causes it, and the four fixes ranked by robustness.ConcurrencyThreads blocked forever on a ReentrantLock with no deadlock reportedWhy an exception between lock() and try leaves a ReentrantLock held forever, how to spot a leaked lock in a thread dump when the JVM reports no deadlock, and the correct acquire shape.ConcurrencyThreadLocal value leaking across requests on a pooled threadWhy a ThreadLocal set on a pool thread outlives the request, how the weak key and strong value actually work, and how the same bug produces both data bleed and a heap leak.ConcurrencyConsumer stuck in Object.wait() with work already in the queueWhy a notify() that arrives before wait() is lost forever, why spurious wakeups are legal, and why the condition must be re-checked in a loop rather than tested once.Concurrencyjava.lang.IllegalMonitorStateException: current thread is not ownerWhy wait(), notify() and unlock() demand ownership, what "current thread is not owner" actually proves about your code, and the two structural bugs that produce it.ConcurrencyThread pool starvation — every worker waiting on a task in its own poolA pooled task that submits to its own pool and waits deadlocks the pool without any lock. How to recognise it in a thread dump and restructure so it cannot happen.ConcurrencyPartially constructed object published by double-checked lockingThe exact reordering that makes non-volatile double-checked locking broken, why it passes thousands of runs before failing, and the three correct lazy-initialisation idioms.ConcurrencyLost update from get-then-put on a ConcurrentHashMapWhy a thread-safe map does not make your read-modify-write thread-safe, why compute() and merge() are atomic while get-then-put is not, and how to pick the right atomic operation.ConcurrencyCompletableFuture failed with nothing loggedWhy a failed CompletableFuture stage produces no log line, how exceptionally, handle and whenComplete differ, and how to make async failures impossible to lose.ConcurrencyawaitTermination never returns and the JVM will not exitWhy shutdown() waits for running tasks, why shutdownNow() cannot stop a socket read, and the shutdown sequence that terminates reliably inside a SIGTERM grace period.ConcurrencyInterruptedException caught and ignored — the task can no longer be cancelledCatching InterruptedException clears the interrupt flag. Why that silently disables cancellation, timeouts and shutdown, and the two correct ways to handle it.ConcurrencyTwo unrelated components sharing a monitor via a boxed Integer or interned StringWhy locking on Integer.valueOf, a String literal, or Boolean shares a monitor JVM-wide, how the autobox cache and string pool cause it, and the safe alternatives.ConcurrencyCache stampede — the same expensive value built many times concurrentlyWhy a cache miss under concurrency triggers a thundering herd of identical loads, why synchronising the whole cache is the wrong fix, and the future-per-key memoiser that works.ConcurrencyLock convoy — throughput collapses as threads are added, with no deadlockWhy a coarse lock on a hot path collapses throughput without deadlocking, how to tell a convoy from a deadlock and from livelock, and how to fix it structurally.ConcurrencyReadWriteLock writer blocked indefinitely behind a stream of readersWhy a non-fair ReentrantReadWriteLock can block a writer indefinitely under continuous reads, why read-to-write upgrade self-deadlocks, and the alternatives that do not starve.ConcurrencyWorker loop never sees the stop flag and runs foreverWhy a worker thread can loop forever on a boolean another thread already set, what the JMM actually guarantees, and why volatile is the fix and sleep() is not.ConcurrencyCommitFailedException: Commit cannot be completed since the group has already rebalancedWhy Kafka revokes your partitions mid-batch, how max.poll.interval.ms and max.poll.records interact, and how to stop the duplicate processing this exception always leaves behind.Distributed systemsConsumer group stuck rebalancing — poll timeout has expiredA consumer group that rebalances forever makes no progress. How generation churn, range assignment, and one slow member create the loop, and how to break it.Distributed systemsThe same message processed twice (at-least-once delivery)Every mainstream broker delivers at least once. Why redelivery is guaranteed rather than exceptional, and the storage-level patterns that make reprocessing harmless.Distributed systemsMessages processed out of order across partitionsOrdering in Kafka is per partition, never per topic. How a key change or a partition-count change reorders events, and the version-guard pattern that makes order irrelevant.Distributed systemsWebhook delivered twice — customer charged twiceWhy payment providers retry webhooks, why a slow handler guarantees duplicates, and how to build a webhook endpoint whose effects are safe under redelivery and reordering.Distributed systemsDatabase and broker diverge after a dual writeWhy writing to a database and a broker in one method can never be atomic, why the compensating-catch fix fails, and how the transactional outbox pattern actually closes the gap.Distributed systemsRetry storm: thundering herd after a dependency failureDiagnose synchronized retries that amplify an outage, then apply bounded exponential backoff, jitter, and retry budgets.Distributed systemsOutbound call has no timeout and exhausts workersTrace a cascading failure from one hung dependency through worker and pool exhaustion, then set deadlines that compose.Distributed systemsDistributed lock lease expired while the holder was still workingWhy a pause lets an old lock holder keep writing, and how lease renewal, fencing tokens, and ownership checks prevent corruption.Distributed systemsClock skew: timestamp ordering or token expiry is inconsistentFind wall-clock assumptions that fail across hosts and replace them with monotonic time, server timestamps, or logical ordering.Distributed systemsExactly-once claim fails at an external side effectSeparate broker exactly-once transactions from external side effects, then build effectively-once outcomes with idempotency.Distributed systemsDead-letter queue growing without an alertTreat a growing DLQ as a correctness incident: identify poison messages, preserve ordering, and replay after fixing the cause.Distributed systemsIdempotency key reused with a different request bodyExplain why an idempotency key must bind to request intent, how to fingerprint canonical bodies, and how to handle legitimate retries.Distributed systemsRead-after-write returned stale data from a replicaDiagnose stale reads after a successful write and use session guarantees, LSN waiting, or primary routing deliberately.Distributed systemsReplication lag: replica served stale data after a writeMeasure write-ahead-log lag, find the slow replay stage, and choose routing or backpressure before stale data becomes corruption.Distributed systems502 Bad Gateway from a reverse proxyDiagnose reverse-proxy 502 responses using upstream timing, direct-origin probes and nginx error logs. Separate connection refusal, invalid headers and crashes.Distributed systems504 Gateway TimeoutFind which proxy timer produced a 504, distinguish connection and read waits, and propagate deadlines so timed-out requests stop consuming backend capacity.Distributed systemsCORS preflight: missing Access-Control-Allow-OriginDiagnose missing Access-Control-Allow-Origin on preflight, authentication failures and cached responses. Fix exact origins, credentials and error-path headers.Distributed systems413 Payload Too LargeTrace a 413 through CDN, nginx and application parsers. Measure encoded body size, scope upload limits and keep memory bounded for large or compressed requests.Distributed systems429 Too Many Requests and Retry-AfterHandle HTTP 429 with delay-seconds or HTTP-date Retry-After, shared quota coordination and a bounded retry budget. Avoid amplifying a rate-limit incident.Distributed systemsnginx 499 client closed requestExplain nginx 499 access-log entries, identify client or outer-proxy deadlines, and stop abandoned requests from holding application and database capacity.Distributed systemsupstream prematurely closed connectionTrace nginx premature upstream closure to crashes, worker timeouts, deployment draining or protocol mistakes. Distinguish header failures from truncated bodies.Distributed systemsIntermittent 502 after an idle keep-alive connectionDiagnose stale upstream connection reuse using idle-gap reproduction and connection timing. Align proxy and origin keep-alive lifetimes without unsafe write retries.Distributed systemsSSL certificate problem: unable to get local issuer certificateUse curl and OpenSSL to separate a missing intermediate certificate from an absent trust root, enterprise TLS interception and the wrong client CA bundle.Distributed systemsERR_INCOMPLETE_CHUNKED_ENCODINGDiagnose truncated HTTP/1.1 chunked responses despite status 200. Compare direct-origin streams, nginx logs and client transfer completion before retrying.Distributed systemsRequest timeouts cascading into pool exhaustionExplain why timed-out HTTP requests still occupy database or connection pools. Diagnose queueing, cancellation and retries, then bound each resource lifetime.Distributed systemsConnection reset by peer on a long-polling endpointFind which hop resets long-poll requests, distinguish active-request idle limits from keep-alive, and design bounded polls with cursors and cancellation.Distributed systemsRedis OOM command not allowed above maxmemorySeparate Redis write rejection from a host OOM kill. Inspect memory policy, TTL eligibility and oversized keys before deciding whether eviction is safe for your data.Distributed systemsMISCONF Redis is configured to save RDB snapshotsFind why Redis background snapshots failed using INFO persistence and server logs. Distinguish storage, permissions and fork pressure without discarding durability policy.Distributed systemsREADONLY You can’t write against a read only replicaIdentify the Redis node actually receiving writes, distinguish reader endpoints from stale failover connections, and restore primary discovery without enabling replica writes.Distributed systemsLOADING Redis is loading the dataset in memoryDetermine whether Redis loading is progressing or restarting. Inspect persistence progress, uptime and storage before tuning readiness checks and bounded client retries.Distributed systemsCROSSSLOT Keys in request don’t hash to the same slotProve slot mismatches with CLUSTER KEYSLOT, choose hash tags around real transaction boundaries, and avoid turning a CROSSSLOT fix into a hot shard or lost atomicity.Distributed systemsMOVED and ASK replies from Redis ClusterHandle Redis Cluster redirects correctly: refresh MOVED slot ownership, pair ASKING with the redirected request on one connection, and validate advertised node addresses.Distributed systemsRedis clients stall during KEYS on a large keyspaceFind blocking keyspace scans with Redis SLOWLOG and command statistics. Replace request-path enumeration with bounded SCAN or an explicit index without assuming snapshot semantics.Distributed systemsA Redis lock lease expires while the worker still runsReconstruct overlapping Redis lease holders, prevent stale unlocks with owner tokens, and enforce correctness at the protected resource when a paused worker resumes.Distributed systemspsycopg2.InterfaceError: connection already closedThe connection object is alive in Python but the socket is gone. How to tell a server-side termination from your own double close, and how pre-ping and recycle actually fix it.PythonRecursionError: maximum recursion depth exceededFind the missing base case, accidental object cycle, or pathological input behind Python RecursionError and fix the algorithm rather than merely raising the limit.PythonMemoryError: unable to allocate arrayDiagnose Python MemoryError by measuring result cardinality and resident memory, then stream, project, chunk, and aggregate data safely.PythonQueuePool limit of size 5 overflow 10 reached, connection timed outQueuePool timeouts mean sessions are not being returned. How to tell a leaked session from genuine saturation, using pool events and pg_stat_activity, and fix the lifecycle.PythonDetachedInstanceError: instance is not bound to a SessionDetachedInstanceError means an ORM object outlived its Session and then needed the database. Why expire_on_commit causes it, and the four correct fixes ranked.PythonRuntimeError: Event loop is closedEvent loop is closed usually means an object outlived the loop that created it, or asyncio.run was called more than once. How to find the owner and fix the lifecycle.PythonTask was destroyed but it is pending!This warning means asyncio work vanished without completing — a garbage-collected task or an abrupt shutdown. Why fire-and-forget tasks lose data, and how to structure them safely.PythonExecuting <Handle ...> took 2.418 seconds (blocked event loop)One synchronous call inside async def stalls every concurrent request. How to measure event-loop lag, find the blocking frame, and move the work off the loop correctly.PythonSettingWithCopyWarning: A value is trying to be set on a copy of a sliceUnderstand pandas view versus copy ambiguity, prove whether an assignment reached the source frame, and make chained indexing deterministic.Pythoncelery.exceptions.WorkerLostError: Worker exited prematurelySeparate Celery worker crashes, SIGKILL from the OOM killer, and hard time limits using worker and kernel evidence.Pythonrequests.exceptions.ReadTimeout: HTTPSConnectionPool read timed outDistinguish DNS/connect/TLS delay from slow response bytes in requests, then set phase-aware timeouts and bounded retries.PythonUnicodeDecodeError: 'utf-8' codec can't decode byteDiagnose byte-versus-text boundaries and identify encodings safely instead of hiding corrupted input with errors=ignore.PythonAssertionError: daemonic processes are not allowed to have childrenWhy a multiprocessing pool worker cannot create a child process, and how to choose threads, a non-daemonic supervisor, or a flat process topology.Python[CRITICAL] WORKER TIMEOUT (pid:1234)Find whether a Gunicorn worker is blocked on CPU, I/O, or a downstream call, then align timeouts and worker models without masking the cause.PythonMutable default argument retains state across callsExplain why Python evaluates defaults once, identify cross-request state leaks, and repair mutable defaults without masking concurrency issues.PythonFATAL ERROR: Reached heap limit Allocation failedDiagnose Node.js heap exhaustion with memory samples and heap snapshots. Separate retained objects, oversized requests and external buffers before changing heap limits.Node.jsUnhandled promise rejection crashes the processFind promises detached from request error handling, understand Node rejection policy and repair async callbacks, finally chains and background task ownership.Node.jsECONNRESET: socket hang up on a reused connectionIdentify ECONNRESET on reused Node HTTP sockets, distinguish idle-close races from upstream failure, and retry only when the operation is safe to repeat.Node.jsMaxListenersExceededWarning: possible EventEmitter memory leakTrace listener registration growth, distinguish intentional fan-out from request leaks, and clean up EventEmitter subscriptions on success, failure and disconnect.Node.jsEADDRINUSE: address already in useDiagnose Node EADDRINUSE using listening sockets and process ancestry. Fix duplicate startup, container namespace confusion and test-server teardown safely.Node.jsERR_HTTP_HEADERS_SENTFind the first response commit behind ERR_HTTP_HEADERS_SENT. Repair missing returns, competing async branches and stream error handling without hiding the race.Node.jsEvent loop blocked by synchronous workMeasure event-loop delay in milliseconds, correlate CPU profiles and request latency, and distinguish synchronous JavaScript, blocking I/O and worker-pool contention.Node.jspg client already connected or released twiceDistinguish pg Client and Pool lifecycles, avoid released-client reuse, and keep transaction queries on one checked-out connection with unconditional cleanup.Node.jsSequelize / Knex pool acquire timeoutFind why Sequelize or Knex cannot acquire a connection: transaction context escapes, leaked borrowers, slow SQL and excess concurrency. Fix ownership before pool size.Node.jsERR_STREAM_PREMATURE_CLOSE during an uploadDiagnose a piped upload closing before completion. Find the first stream failure, distinguish client aborts from storage errors and publish files only after validation.Node.jsProcess exits before asynchronous writes finishUnderstand why process.exit truncates output, why promises alone do not keep Node alive, and how to drain servers, streams and database pools before shutdown.Node.jsERR_MODULE_NOT_FOUND during ESM/CommonJS migrationDiagnose missing ESM imports using the emitted file, explicit extensions, package exports and production dependencies. Separate resolution failures from interop errors.Node.jsUnbounded JSON body blocks the loop or exhausts memoryPrevent oversized JSON requests from exhausting Node memory or blocking the event loop. Enforce byte limits before buffering, account for decompression and bound concurrency.Node.jsUndici / fetch connections remain occupiedDiagnose fetch queues and socket pressure caused by unread response bodies, per-request dispatchers and unlimited concurrency. Fix body ownership and dispatcher lifecycle.Node.jsERR_UNHANDLED_REJECTION in a worker threadRepair unhandled async worker callbacks, distinguish task failures from worker crashes, and settle parent promises once when a worker errors or exits.Node.jscommand not found in a script that works interactivelyTrace why cron, systemd or CI cannot find a command that works in your terminal. Separate PATH lookup, shell functions, runtime activation and working-directory errors.Linux / shellPermission denied when executing a scriptDiagnose execution permission failures using namei, findmnt and the real service identity. Distinguish missing execute bits, inaccessible interpreters and noexec mounts.Linux / shellbad interpreter: No such file or directory with CRLFExpose the carriage return hidden in a Linux shebang, distinguish it from a missing interpreter or ELF loader, and repair line endings without changing script behaviour.Linux / shellArgument list too longFix Linux argument-list limits without losing filenames containing spaces or newlines. Understand glob expansion, environment size and why xargs with a glob still fails.Linux / shellToo many open filesMeasure the failing process’s descriptor count and limits, separate leaks from bounded concurrency, and fix file or socket ownership before changing LimitNOFILE.Linux / shellOut of memory: Killed processInvestigate a SIGKILL using kernel logs and cgroup v2 memory.events. Separate container limits, host pressure and application allocation failures before sizing memory.Linux / shellNo space left on device despite free disk spaceUse df -i on the actual target filesystem to distinguish inode exhaustion from full blocks, wrong mounts and deleted open files. Fix small-file growth at its source.Linux / shellText file busy during executable replacementUnderstand ETXTBSY when overwriting a running executable or executing a file still open for writing. Use immutable release artifacts and same-filesystem rename.Linux / shellset -e script continues after a failed pipelineReproduce why Bash errexit misses upstream pipeline failures, conditional functions and command substitutions. Use pipefail plus explicit checks at publication boundaries.Linux / shellAn unquoted variable turns one argument into severalSee the exact argument vector produced by Bash expansion. Preserve paths with quotes, use arrays for command arguments, and handle empty values and leading dashes explicitly.Linux / shellOOMKilled — container exit code 137Distinguish a container memory-limit kill from node pressure and other SIGKILL exits. Inspect termination state, cgroup counters and peak allocation before tuning limits.Linux / shellCrashLoopBackOffTrace CrashLoopBackOff to the previous container exit, probe failure or completed foreground process. Preserve evidence and repair the cause of repeated restarts.Linux / shellImagePullBackOff / ErrImagePullRead the registry failure behind ImagePullBackOff. Separate missing images, pull-secret scope, node egress, rate limits and incompatible image platforms.Linux / shellReadiness probe failedCompare the kubelet probe with the request that succeeds. Check pod IP binding, port, HTTP contract, timeouts and dependency checks before changing readiness.Linux / shellLiveness probe failed during startupProve that kubelet interrupts legitimate initialisation, separate startup from liveness, and size a startup probe using cold-start measurements under real limits.Linux / shellCreateContainerConfigError — missing Secret or ConfigMapFind the exact unresolved configuration reference before container startup. Check namespace, key names, generated resource names and secret-controller reconciliation.Linux / shellEvicted — node disk pressureDistinguish node DiskPressure from a pod storage-limit eviction. Trace logs, writable layers, emptyDir and inode consumption without deleting runtime state.Linux / shellrunAsNonRoot and image will run as rootResolve Kubernetes runAsNonRoot validation with an explicit numeric UID and compatible file ownership. Distinguish UID validation from application permission failures.Linux / shellJVM container killed despite available Java heapUnderstand why MaxRAMPercentage and Xmx do not cap total Java process memory. Compare cgroup usage, effective JVM flags, thread stacks and native allocation.Linux / shellexec format error — container architecture mismatchTrace exec format error through node architecture, image manifests and the actual executable. Fix cross-builds and distinguish ELF mismatch from broken entrypoint scripts.Linux / shellno space left on device — container layersDiagnose ENOSPC in image pulls, builds and writable layers. Understand OverlayFS copy-up, immutable lower layers, inodes and Docker daemon storage limits.Linux / shellSIGTERM misses the application — shutdown ends in SIGKILLTrace shutdown signals through shell wrappers, PID 1 and application handlers. Use exec correctly, budget preStop time and test draining with work in flight.Linux / shell