Node.js Worker Threads for CPU-Intensive Tasks
Your API handles image resizing inline, and under load every other request stalls while one big image gets processed. Node.js is single-threaded for JavaScript execution — a CPU-heavy synchronous task (image processing, large JSON parsing, cryptographic hashing, complex calculations) blocks the event loop, and every other request, websocket message, and timer waits behind it. Worker threads exist to move that work off the main thread entirely.
Why async/await doesn't help here
This is the confusion that trips people up first: async/await solves I/O concurrency (waiting on network requests, file reads, database queries) by letting Node do other work while waiting. It does nothing for CPU-bound work — a synchronous for loop crunching numbers blocks the thread regardless of how many async functions wrap it, because there's no I/O wait to yield during.
// This still blocks the event loop — async doesn't parallelize CPU work
async function hashPassword(password) {
return crypto.pbkdf2Sync(password, salt, 100000, 64, "sha512");
}
Moving work to a worker thread
// worker.js
const { parentPort, workerData } = require("worker_threads");
function heavyComputation(n) {
let result = 0;
for (let i = 0; i < n; i++) result += Math.sqrt(i);
return result;
}
parentPort.postMessage(heavyComputation(workerData.n));
// main.js
const { Worker } = require("worker_threads");
function runWorker(n) {
return new Promise((resolve, reject) => {
const worker = new Worker("./worker.js", { workerData: { n } });
worker.on("message", resolve);
worker.on("error", reject);
});
}
app.get("/compute", async (req, res) => {
const result = await runWorker(1_000_000_000);
res.json({ result });
});
The heavy loop now runs on a separate OS thread with its own V8 instance. The main thread's event loop stays free to handle other requests while it waits.
Sharing data efficiently
Workers communicate by message-passing by default, which involves structured-clone serialization — fine for small payloads, expensive for large ones. For big buffers (image data, large arrays), use SharedArrayBuffer or transfer ownership instead of copying:
const worker = new Worker("./worker.js", {
workerData: { buffer: largeArrayBuffer },
transferList: [largeArrayBuffer], // ownership moves, no copy
});
After transferring, the main thread's reference to that buffer becomes unusable — ownership genuinely moved, it wasn't cloned.
A worker pool for repeated work
Spawning a new worker per request has real overhead — thread creation isn't free. For frequent CPU-bound tasks, use a pool:
const { Piscina } = require("piscina");
const pool = new Piscina({ filename: "./worker.js" });
app.get("/compute", async (req, res) => {
const result = await pool.run({ n: 1_000_000_000 });
res.json({ result });
});
Piscina (or Node's own worker_threads combined with a hand-rolled pool) reuses a fixed set of worker threads instead of paying startup cost on every request.
Common mistakes
- Using worker threads for I/O-bound work (database calls, HTTP requests) — that's what
async/awaitalready solves efficiently; workers add overhead for no benefit there. - Spawning a new worker per request under real load without pooling, causing thread churn that can be slower than just blocking the main thread for lightweight tasks.
- Passing huge objects through
postMessagewithouttransferList, silently paying a full structured-clone cost on every message. - Assuming worker threads share memory like threads in other languages by default — they don't unless you explicitly use
SharedArrayBuffer; regular objects are copied, not shared.
What I'd actually use
Worker threads (via a pool like Piscina) for genuinely CPU-bound work that's frequent enough to justify the setup — image processing, PDF generation, complex data transformation. For truly heavy, infrequent batch work, a separate background job service (queued via Redis/BullMQ) is often the better architectural fit than keeping it inline in your API process at all.
Next steps
Profile your API under load (node --prof or a APM tool) and look for request handlers with long synchronous CPU time — that's your worker-thread candidate list, not everything that "feels slow," since a lot of perceived slowness is actually I/O latency that workers won't fix.