What you'll learn
Quick Answer
Concurrency is about structure: splitting a program into independent tasks that can be in progress at the same time. Parallelism is about execution: actually running more than one of them at the same instant, which needs multiple CPU cores. A single-threaded Node process is concurrent but not parallel. Use concurrency for I/O-bound work and parallelism for CPU-bound work.
Structure versus execution
Rob Pike put it best: concurrency is dealing with many things at once; parallelism is doing many things at once. One is about how you structure a program, the other about how it actually runs.
A concurrent program is split into tasks that can make progress independently. Whether they truly run side by side depends on the hardware. On a single core the CPU interleaves them: task A runs, pauses at a wait, task B runs, then control returns to A. All are in progress, but only one executes at any instant.
Parallelism is the stronger claim: two tasks executing at the same physical moment on two cores. Every parallel program is concurrent, but plenty of concurrent programs are never parallel, and often do not need to be.
The useful question is never which is better. It is whether your workload spends its time waiting on I/O or burning CPU. That answer decides which model helps.
Concurrency on one thread
Node runs your JavaScript on a single thread with an event loop. When you await a network call the thread does not block: it parks that task, picks up other work, and resumes when the response lands. That is concurrency with no parallelism.
Three fake API calls of 300ms each, sequentially versus overlapped:
// sequential: await one, then the next
await fakeApiCall('A', 300);
await fakeApiCall('B', 300);
await fakeApiCall('C', 300);
// 916 ms (the waits add up)
// concurrent: start all, then await together
await Promise.all([
fakeApiCall('A', 300),
fakeApiCall('B', 300),
fakeApiCall('C', 300),
]);
// 303 ms (the waits overlap)Same single thread, same three calls. The concurrent version finishes in roughly the time of the longest call rather than the sum. Nothing ran in parallel; the thread simply stopped sitting idle during each wait.
The event loop is what makes this work. Every time your code hits an await, the function suspends and the loop is free to run timers, accept new connections, or resume a different suspended function whose data has arrived. One Node process serves thousands of simultaneous connections this way, because at any instant almost all of them are only waiting.
Parallelism needs cores
Concurrency does nothing for work that never waits. A loop counting prime numbers has no idle moment to fill; the CPU is the bottleneck. To speed it up you must run parts of it on different cores at the same time.
Node exposes real OS threads through worker_threads. Counting primes from 0 to 3,000,000, split across four workers:
single thread: 2413 ms (216816 primes)
4 workers: 616 ms (216816 primes)
speedup: 3.92xThat near four times drop is parallelism: four cores doing arithmetic simultaneously. Workers do not share memory by default; each has its own heap and you pass data by messages or a SharedArrayBuffer. That isolation is the price, and it is why parallelism takes more setup than a single await.
Spawning a worker is not free either: it starts a fresh V8 instance and costs a few milliseconds. For repeated jobs, keep a pool of workers alive and feed them tasks rather than creating one per job. The child_process and cluster modules are the other routes to multiple cores, trading shared memory for full process isolation.
The GIL and single-thread runtimes
Language runtimes make different choices here.
- Node.js runs one thread for your JS by design. Concurrency is free and idiomatic; parallelism means spawning workers or child processes.
- CPython has a Global Interpreter Lock. Even with the
threadingmodule, only one thread runs Python bytecode at a time, so threads give you I/O concurrency but no CPU speedup. For that you usemultiprocessing, which forks separate interpreters. CPython 3.13 ships an experimental free-threaded build that removes the GIL. - Go, Java, Rust, C# run threads in parallel across cores with no interpreter lock.
A concrete case: a Python script fetching 100 URLs finishes far faster with a ThreadPoolExecutor, because each thread sits blocked on the network and the GIL is released during that wait. The same script doing 100 CPU-heavy hashes gets no speedup from threads at all; only ProcessPoolExecutor helps. asyncio is Python's event-loop answer for the I/O case, mirroring Node.
The pattern: interpreted, garbage-collected runtimes often serialise execution for safety, then offer a separate escape hatch for true parallelism.
Choosing the model
Decide by asking what the work spends its time doing.
- I/O-bound work, such as database queries, HTTP calls, file reads, or waiting on a queue, wants concurrency. One thread with async/await handles thousands of in-flight operations because each is mostly idle. Extra cores buy almost nothing.
- CPU-bound work, such as image resizing, parsing, encryption, or number crunching, wants parallelism. Split it across cores with workers or processes. Async here is pointless because there is no wait to overlap.
The gotcha: put a CPU-heavy loop directly inside an async request handler and it blocks the event loop. Every other in-flight request freezes until it finishes, because they all share that one thread. Running the 2.4-second prime count inline would stall the entire server for 2.4 seconds.
Real systems use both. A web server is concurrent by nature, and when a request needs something heavy, such as generating a PDF or resizing an upload, it hands that one task to a worker pool and awaits the result. The handler stays responsive while the CPU work runs off the event loop. Measure first: profile where wall-clock time actually goes, because if it is I/O no amount of threads will help, and if it is CPU no amount of async will.
