The real cost of database connections in AWS Lambda
A customer brought me in to review their AWS setup. The mandate was to cut costs and improve performance. One of the first bottlenecks we uncovered was not in the code itself but in how Lambda interacted with the database. At first glance it looked like the system was sized correctly, but closer inspection showed that resources were being wasted just on connection handling rather than on useful query work.
Where the time went
Every Lambda invocation created a new database connection, ran a query, and then closed the session. The queries were simple and efficient, but the connection lifecycle was anything but. Establishing a TCP session, negotiating TLS, and authenticating each time created a consistent delay. The function’s execution time was being dominated by the mechanics of opening a channel rather than by the query itself. When functions are short-lived and run hundreds of times per second, that overhead quickly adds up.
The problem was not only about latency. Databases are designed to handle a steady pool of long-lived connections, not thousands of rapid, spiky connection attempts. Under bursts of Lambda activity the database was spending a large share of CPU and memory just managing handshakes, socket buffers, and authentication checks. Query throughput was modest, but the infrastructure was behaving as though it was under attack. To keep the system stable the customer had provisioned a larger RDS instance than their actual workload required. They were paying for capacity that existed only to absorb connection churn.
The quick hack
One workaround is to move the database client into the global scope of the Lambda function so that warm starts can reuse the same connection. This trick helps when AWS reuses the same execution environment across multiple invocations, because the connection remains open in memory and the function can skip the handshake. Latency improves and the database sees fewer new connection attempts.
The drawback is that environment reuse is not guaranteed. When traffic ramps up, AWS spins up new execution environments to handle the load, and each of them still opens its own connections. With high concurrency, dozens or hundreds of parallel environments all keep their own sockets open. The result is a database flooded with idle or half-used connections that have to be tracked and maintained. The global scope hack reduces pressure in light usage, but it is fragile under bursty or production-scale workloads.
The right fix
The sustainable solution is to use RDS Proxy. The proxy sits between Lambda and the database and maintains a pool of backend connections. Lambda functions connect to the proxy endpoint, which is fast and lightweight, and the proxy handles the actual session reuse with the database. From the database’s perspective, it is dealing with a stable and predictable set of long-lived sessions, not a storm of short-lived handshakes. From Lambda’s perspective, the connection process becomes nearly instantaneous, and execution time is dominated again by the query rather than the setup.
Integrating the proxy requires some network and authentication alignment but not much else. Lambda and the proxy run in the same VPC, the client points to the proxy endpoint instead of the database endpoint, and credentials are managed by Secrets Manager or IAM. The Lambda code does not need to know how pooling works; the proxy absorbs the churn. What matters is that session reuse becomes systematic and reliable, removing the dependency on environment reuse and insulating the database from traffic spikes. In this customer’s case, the result was both faster queries and the ability to shrink the RDS instance size, since capacity no longer had to be reserved for connection overhead.
Takeaway
Serverless and relational databases have a fundamental mismatch. Lambda is designed to spin up and down with demand, but relational engines are tuned for a smaller number of stable connections. Without a pooling layer the friction shows up as latency and inflated costs. Tricks like global scope reuse can buy time during development, but production workloads demand something more robust.
The customer I worked with learned this first-hand. Their performance issues had little to do with query logic and everything to do with connection management. Once RDS Proxy was in place, the database workload stabilized, execution times dropped, and the monthly bill decreased. The lesson is clear: before scaling Lambda with a database backend, make sure connection management is solved, because otherwise you will be paying for wasted cycles rather than real work.
Member discussion