Java / JVM

HikariPool-1 — Connection is not available, request timed out after 30000ms

Written and reviewed by Sahil Srivastav

HikariCPConnection poolProduction incident
java.sql.SQLTransientConnectionException: HikariPool-1 - Connection is not available, request timed out after 30000ms.
	at com.zaxxer.hikari.pool.HikariPool.createTimeoutException(HikariPool.java:696)
	at com.zaxxer.hikari.pool.HikariPool.getConnection(HikariPool.java:197)

What this error actually means

This is a queue timeout, not a database error. A thread asked the pool for a connection, the pool had none free, and the thread waited `connectionTimeout` milliseconds (30 seconds by default) before giving up. The database itself may be completely idle and healthy while this exception floods your logs.

The distinction that matters: the pool is a fixed set of slots, and a slot is only returned when your code closes the connection. So there are exactly two possible states behind this message — either every slot is legitimately busy doing work (saturation), or slots have been borrowed and never given back (a leak). Those two states need opposite fixes, which is why raising the pool size blindly resolves one and makes the other worse.

A leak produces a characteristic signature: the failure appears some number of requests *after* the request that actually caused it, and it never recovers until the process restarts. Saturation, by contrast, tracks traffic — it appears at peak and clears when load drops.

Causes, most common first

  1. 1A connection borrowed on a path that throws. The happy path closes the connection; the exception path returns early and skips the close. Every failed request permanently consumes one slot. This is the single most common cause, and the one most often misdiagnosed as "we need a bigger pool".
  2. 2A transaction left open before the connection is returned. The connection is closed but neither committed nor rolled back. With autocommit off, the session goes back to the pool still holding a snapshot and possibly row locks — PostgreSQL reports it as `idle in transaction`. The slot counts as available while being useless.
  3. 3Genuine saturation: pool smaller than concurrent demand. Request concurrency times average query time exceeds pool size times the timeout window. Common after adding a slow report endpoint, or when the pool was sized for one instance and the service was scaled to ten against the same database `max_connections`.
  4. 4Long-running work holding a connection it does not need. A handler fetches rows, then makes an HTTP call, then writes — all inside one borrowed connection. The pool slot is held for the duration of the network call. Under any latency spike in the downstream service, the pool drains.
  5. 5Nested acquisition from a pooled thread. A task holding a connection submits sub-tasks that also need connections, from a bounded executor. Classic starvation deadlock: the holder waits for the sub-tasks, the sub-tasks wait for the pool, and the pool waits for the holder.

When you see it

  • Errors begin minutes after a burst of failed or exceptional requests, not during it
  • A restart fixes it completely, then it reappears after roughly the same amount of traffic
  • Database CPU is low and `pg_stat_activity` shows sessions sitting in `idle in transaction`
  • Active connections plateau at exactly `maximumPoolSize` and never drop
  • Latency on unrelated endpoints degrades first, because they queue behind the same pool

How to diagnose it

Step 1

Turn on leak detection first

HikariCP will log a stack trace of the exact borrow site that has held a connection too long. This converts guesswork into a file and line number. Set it slightly above your slowest legitimate query.

spring.datasource.hikari.leak-detection-threshold=20000

Step 2

Ask the database what the sessions are doing

This is the decisive query. Many rows in `idle in transaction` means a lifecycle bug in your code. Many rows in `active` with long durations means real saturation or missing indexes.

SELECT state, count(*), max(now() - state_change) AS longest
FROM pg_stat_activity
WHERE datname = current_database()
GROUP BY state ORDER BY 2 DESC;

Step 3

Check whether the pool ever recovers

Expose the HikariCP metrics and watch `hikaricp.connections.active` after traffic stops. If it stays pinned at maximum with zero requests in flight, you have a leak — no amount of tuning will help.

curl -s localhost:8080/actuator/metrics/hikaricp.connections.active

Step 4

Reproduce it deliberately

Fire N requests at the endpoint you suspect, where N is your pool size, and force each one to fail. If the next healthy request times out, you have proven the exception path is the leak.

The fix

For the leak, make release unconditional. Acquire the connection inside try-with-resources so the close happens on every path, including the ones you did not anticipate. Never call `close()` at the end of a method body and hope no exception fires before it.

For the open transaction, make the transaction boundary explicit and finish it before the connection is released: commit on success, roll back on failure, and only then close. If you are on a framework, the equivalent is to stop manually acquiring the connection at all and let the declarative transaction manager own the whole boundary.

For genuine saturation, stop holding connections across work that does not need the database. Move HTTP calls and heavy in-memory transformation outside the borrow window, then size the pool from measurement rather than intuition: a pool of roughly `(cores * 2) + effective_spindles` per instance, multiplied by instance count, must stay under the database `max_connections` with headroom for migrations and admin sessions. Bigger is very often slower.

For nested acquisition, break the dependency: either one connection travels down with the sub-tasks, or the sub-tasks use a separate pool, or the work is restructured so a holder never waits on a borrower.

// Leaks on any throw between getConnection() and close()
Connection c = dataSource.getConnection();
var rs = c.createStatement().executeQuery(sql);   // throws -> slot gone forever
c.close();

// Release is unconditional, and the transaction ends before the slot returns
try (Connection c = dataSource.getConnection()) {
    c.setAutoCommit(false);
    try (var st = c.prepareStatement(sql)) {
        st.setLong(1, customerId);
        try (var rs = st.executeQuery()) {
            // ... map rows
        }
    }
    c.commit();
} catch (SQLException e) {
    // rollback happens on the same connection before close(); see note below
    throw new DataAccessException(e);
}

How to stop it coming back

  • Keep `leakDetectionThreshold` enabled in staging permanently — it costs nothing and catches the bug before production does
  • Alarm on `hikaricp.connections.pending > 0` sustained, not on the timeout exception; pending is the early signal and the exception is the aftermath
  • Add a test that forces the exception path N+1 times and asserts the pool still serves traffic. Lifecycle bugs are invisible to happy-path tests
  • Cap statement time at the database (`SET statement_timeout`) so one pathological query cannot hold a slot indefinitely
  • Review any code that acquires a connection outside the framework’s transaction management — that is where these bugs live

Practise this failure in a real repository

Gronex ships this exact failure as a runnable repository: a small pool, an exception path that skips cleanup, and sessions returned while still idle in transaction. The bundled test suite asserts the lifecycle invariant rather than the happy path, so a bigger pool size will not make it pass.

FAQ

Should I just increase maximumPoolSize?

Only after `pg_stat_activity` shows the sessions genuinely active. If they are idle in transaction, a bigger pool buys you a longer runway before the same outage and adds load the database has to track. Fix the lifecycle first, size second.

Why does the error appear long after the request that caused it?

Because a leak consumes one slot per failure and the pool keeps serving traffic from remaining slots. You only see the timeout once the last slot is gone, which can be hundreds of requests later — this lag is exactly why the wrong endpoint usually gets blamed.

Does closing the connection roll back an open transaction?

Do not rely on it. Behaviour depends on the driver and pool configuration, and HikariCP may reset state on return, silently discarding work you believed was committed. Make the commit or rollback explicit so the outcome is in your code, not in configuration.

Is this the same as "too many clients already"?

No, and the difference is diagnostic. This message comes from your application pool refusing to wait longer. `FATAL: sorry, too many clients already` comes from PostgreSQL refusing a new connection because `max_connections` is reached — usually the result of too many instances each with a generous pool.

Related

Other errors engineers hit next to this one

Full error and symptom index →