← All posts

Node.js Memory Leak Debugging: A Practical Guide

By Harshit Dixit · · 12 min read

A Node.js memory leak rarely announces itself. The process just grows, a little with every request, until the container is killed for exceeding its memory limit or V8 gives up with FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory. It restarts, everything looks fine, and a few hours later it happens again.

This guide is the process that reliably finds the cause: confirm it's actually a leak, capture heap snapshots without taking production down, compare them properly, and fix the handful of patterns that cause most leaks.

Step 1: confirm it's a leak, not just high memory

High memory use isn't a leak. V8 grows the heap lazily and doesn't hand memory back to the operating system eagerly, so a process that sits at 600 MB can be perfectly healthy. A leak is memory that keeps growing after garbage collection under a steady workload.

Log the numbers that matter over time:

setInterval(() => {
  const { rss, heapUsed, heapTotal, external, arrayBuffers } = process.memoryUsage();
  const mb = (n) => Math.round(n / 1024 / 1024);
  console.log(JSON.stringify({
    rss: mb(rss),
    heapUsed: mb(heapUsed),
    heapTotal: mb(heapTotal),
    external: mb(external),
    arrayBuffers: mb(arrayBuffers),
  }));
}, 60_000).unref();

Then read the shape of the graph:

  • heapUsed saw-tooths but the troughs keep rising— classic JavaScript heap leak. Objects are surviving garbage collection that shouldn't. This is what heap snapshots find.
  • rss grows but heapUsed is flat — the growth is outside the JS heap: buffers, native addons, or memory fragmentation. Check external and arrayBuffers; heap snapshots won't show this.
  • Memory climbs under load and falls back when traffic drops — not a leak, just concurrency. Each in-flight request holds memory. Fix it with limits, not a leak hunt.

Also check that the heap limit matches the container. V8 picks a default heap size that doesn't always track the container's memory limit, so set it explicitly and leave headroom for everything outside the heap:

# a 1 GB container: give the JS heap ~75% of it
node --max-old-space-size=768 server.js

Step 2: reproduce it locally

Leaks are far easier to find when you can trigger them on demand. Run the app locally with the inspector enabled and drive the suspect endpoint with a load tool:

node --inspect server.js

# in another terminal: steady load against the suspect route
npx autocannon -c 20 -d 120 http://localhost:3000/api/orders

Open chrome://inspect in Chrome, click inspect under your Node process, and go to the Memorytab. If you don't know which endpoint leaks, the memory log from step 1 plus your request logs will usually narrow it down — the growth correlates with a route or a job.

Step 3: the three-snapshot technique

A single heap snapshot is nearly useless: it shows everything that exists, most of it legitimately. What you want is what's accumulating. The reliable method:

  1. Warm the app up (a few hundred requests), so caches and lazy modules are loaded.
  2. Take snapshot 1. (Taking a snapshot forces a full garbage collection first.)
  3. Run a fixed amount of load — say, 1,000 requests to the suspect route.
  4. Take snapshot 2.
  5. Run the same load again.
  6. Take snapshot 3.

Select snapshot 3, switch the view to Comparison, and compare against snapshot 2. Sort by # Delta or Size Delta. Object types that grew by roughly the same amount between 1→2 and 2→3 — and in proportion to the requests you sent — are your leak. One-off growth between 1 and 2 that doesn't repeat is usually just warm-up.

Reading the retainers

Finding the leaking objects is half the job; the other half is finding what's keeping them alive. Click a leaking object and look at the Retainers pane underneath. It shows the chain of references from a garbage-collection root down to that object. Read it bottom-up until you reach something you recognise from your own code — a module-level Map, an event emitter, a closure in a particular file. That's the thing holding on.

Names help enormously here. Anonymous arrow functions and plain object literals show up as (closure) and Object; named functions and class instances show up by name. If the snapshot is unreadable, temporarily wrapping suspect data in a named class is a legitimate debugging trick.

Step 4: capturing snapshots in production

Sometimes a leak only happens with production traffic. You can capture a snapshot from a running process, but understand the cost first: writing a snapshot pauses the process — for seconds on a large heap — and needs substantial extra memory while it runs. Take the instance out of the load balancer first, or do it on one replica you can afford to lose.

Three ways, from simplest:

# 1. Snapshot on a signal, no code changes:
node --heapsnapshot-signal=SIGUSR2 server.js
kill -USR2 <pid>          # writes Heap.<date>.<pid>.heapsnapshot to the cwd

# 2. Snapshot automatically as the heap nears its limit:
node --heapsnapshot-near-heap-limit=2 server.js

# 3. From code, e.g. behind an authenticated admin endpoint:
import { writeHeapSnapshot } from "node:v8";
const file = writeHeapSnapshot(); // returns the file path

Copy two or three snapshots taken some time apart to your machine and load them into DevTools' Memory tab — the same comparison works on files.

The usual suspects

Most Node.js leaks are one of these.

1. Unbounded caches

The most common leak by far: a module-level Map or object used as a cache, keyed by something with unbounded variety — user IDs, URLs, query strings — with nothing ever evicting entries.

// leaks: one entry per distinct URL, forever
const cache = new Map();
export async function getPage(url) {
  if (!cache.has(url)) cache.set(url, await render(url));
  return cache.get(url);
}

Every in-process cache needs a bound — a maximum size, a TTL, or both. Use a proper LRU (the lru-cache package is the standard choice) or move the cache to Redis, where memory is managed and shared across instances.

2. Event listeners that are never removed

Adding a listener to a long-lived emitter inside a request handler adds one listener per request. Node warns you with MaxListenersExceededWarning: Possible EventEmitter memory leak detected — take that warning seriously rather than raising the limit.

// leaks: a new listener on a global emitter for every request
app.get("/stream", (req, res) => {
  bus.on("update", (data) => res.write(data));
});

// fixed: remove it when the request ends
app.get("/stream", (req, res) => {
  const onUpdate = (data) => res.write(data);
  bus.on("update", onUpdate);
  req.on("close", () => bus.off("update", onUpdate));
});

An AbortController makes this tidier when you have several listeners: pass its signal to events.on or addEventListener and abort once on cleanup.

3. Timers and intervals

A setInterval keeps its callback — and everything the callback closes over — alive until it is cleared. Intervals created per connection or per job and never cleared are a steady leak. Always keep the handle and clearInterval it on cleanup.

4. Closures capturing more than they need

A small callback that's stored somewhere long-lived keeps its whole enclosing scope reachable. If that scope holds a large request body or a parsed file, it all stays in memory. Copy out just the values the callback needs.

5. Promises that never settle

A promise waiting on something that never happens — a response from a socket that died, a queue message that was dropped — holds its callbacks and their closures forever. Put timeouts on every external wait; AbortSignal.timeout(ms) works with fetch and many other APIs.

6. Growing arrays for “later”

Logs, metrics or audit events pushed into an in-memory array to be flushed later, where the flush fails or never runs. Cap the buffer, and drop or flush when it's full.

When to reach for WeakMap and WeakRef

If you need to associate data with an object you don't own — metadata about a request object, say — a WeakMaplets the entry disappear when the key object is garbage-collected. It's the right tool for that specific case. It is nota general cache: keys must be objects, and you can't control or predict when entries go. For a cache, use an LRU.

A checklist

  1. Log process.memoryUsage() and confirm the post-GC baseline is rising.
  2. Decide if it's JS heap (heapUsed) or outside it (rss, external).
  3. Set --max-old-space-size to fit the container.
  4. Reproduce locally with --inspect and a load tool.
  5. Take three snapshots with identical load between them; compare 3 against 2.
  6. Follow the retainers of the objects that grow per request.
  7. Check the usual suspects: caches, listeners, timers, closures, pending promises, buffers.
  8. Fix, re-run the same load, and confirm the baseline is flat.

If a leak is taking down production and you need a second pair of hands on it, that's the kind of problem a freelance senior Node.js engineer can be brought in for.

// free, no strings

Book a free consultation

30 minutes, no cost, no obligation — pick a slot that works for you.

  • ›Sign in with Google to book instantly
  • ›Or just email/WhatsApp if you'd rather
  • ›We'll follow up with a free draft, not a sales pitch