Skip to content

Multi-threading

Moving a task into a worker still leaves coordination work on the host. Only the task body leaves the request thread: the host still accepts the request, routes it, encodes the arguments, submits the call, receives the result, and writes the response. All of that costs CPU, so pick a worker count that leaves room for it.

WorkloadStarting pointWhy
HTTP service with a few CPU-heavy routesthreads: 1Keeps expensive work off the event loop without taking many cores from the host.
HTTP service with a growing queue of heavy callsthreads: 2 to 4More workers can improve tail latency while requests wait for CPU.
Independent CPU-bound batch jobsavailable cores - 1The worker bodies can execute in parallel; leave a core for the host and OS.
Mostly network, database, or filesystem workKeep it async; do not add workers just for I/OA worker does not make an external dependency faster.

Treat these as places to start. Task duration, payload size, request rate, and the runtime itself all affect the right worker count.

Move the expensive route work into tasks and leave the rest of the handler on the request thread. One worker is usually enough to clear the head-of-line blocking, and it keeps most of the machine available to the host.

import { createPool } from "knitting";
const threads = 1;
const handlers = createPool({
threads,
})({ renderSsrPage, issueJwt });

The Hono example is built this way: SSR and JWT run in the worker, /ping stays on the request thread. Under a saturating mixed load, /ping served 201% more requests than the single-threaded version, because it was no longer stuck behind a render. See the Hono server example and its 16-core measurements.

A second worker helps when heavy calls are queueing behind the first one and the host still has CPU to spare. Each worker you add competes with the host for cores and brings its own scheduling, memory, and result draining.

In the Hono workload, one worker gave the highest throughput under saturation. Larger pools did cut the p99 of the heavy routes at a fixed offered rate, but they could not make the host a faster producer, and that is what a saturated run measures. This is the usual outcome when a single event loop feeds the pool.

In compatible pools with more than one worker, native work stealing is automatic. Workers pull from shared work instead of each sitting on a private backlog, so whichever one is free takes the next task. That helps balance tasks with different durations. The host-side cost stays where it is.

What a worker does when it runs out of work

Section titled “What a worker does when it runs out of work”

Idle CPU is a resource the host can use, so a worker that has caught up spends as little of it as it can, in this order:

  1. Finished results go out first. In a stealing pool, a worker sends its response before it tries to claim more work.
  2. Safe return-side releases are drained. A BufferReference is released only once its result has been consumed or released explicitly.
  3. The worker looks for available shared work. If there is none, it makes a best-effort maybeGc() call where the runtime offers one, then parks.

How it parks depends on the size of the pool:

  • A single worker spins for 50 µs first, because it sits on the request’s critical path.
  • Multi-worker pools skip the spin and park immediately, leaving the CPU to the host or to an awake peer.

Both are defaults. worker.timers is there if you need to override them.

The host has a separate wait path for completed results. host.doorbell is enabled by default when the runtime can wake the host without polling:

Runtime or topologyCompletion path
Node and Bun thread workersAtomics.waitAsync
Deno thread workersA thread-safe FFI callback when FFI permission is available
Process workersA process-local completion transport
Compiled workers and browsersPolling/fallback; host options are unsupported for compiled workers

If a runtime denies the required capability, Knitting falls back to portable polling. Set host: { doorbell: false } when you need a controlled polling comparison or when workers already occupy every core. Node thread workers can also request host.nativeDoorbell: true, which uses the optional native uv_async_t bridge from the knitting_doorbell addon. It is off by default, requires the regular doorbell, and does not apply to process workers.

Hono mixed load · less is more

Leave more CPU for serving requests.

Each bar is CPU time for one served request. Multi-worker pools share available work through native stealing, then sleep when there is none.

Host 8,208 rps
2 workers 15,811 rps
4 workers 13,906 rps
8 workers 11,649 rps
15 workers 11,987 rps

Full method and measurements

Each bar is total server CPU divided by completed requests under saturating mixed Hono load, so a shorter bar leaves more CPU for the host and for everything else on the machine. One worker is the efficient point for this workload: 100 CPU-µs per request at 18,014 RPS. A larger pool can still be useful if lower tail latency is worth the extra CPU cost.

In these measurements, pools with more than one worker are not saturated. The host is the only producer, so the extra workers spend most of their time waiting for a call that may never arrive. A pool that spun through that wait would look busy without being useful: at fifteen workers, the old 750 µs spin budget burned 6.4 idle cores and 954 CPU-µs per request to serve fewer requests than the current policy serves at 236. Parking idle workers keeps that CPU cost down.

Measure each candidate thread count under two types of load:

  1. Saturating mixed load shows the ceiling, and whether cheap endpoints are still delayed by expensive ones.
  2. Fixed offered rate makes p50 and p99 comparable, since every candidate receives the same work.

Record server CPU, idle CPU, throughput, and p50/p99 for each route. Total RPS on its own will mislead you: a configuration that cuts a heavy route’s tail latency while giving up some throughput is often the one you want.

For the lower-level options, see Creating pools, Performance, and Work stealing.