Performance
How to read the ratings
Section titled “How to read the ratings”These ratings are a rough guide to the cost of one call. A Slow entry can still be 2–4× faster than
postMessage, depending on the payload and the work being done.
- Best — almost no per-call overhead
- Fast — inexpensive for most workloads
- Good — suitable for normal workloads
- Fair — fine occasionally; watch hot paths
- Slow — consider another representation in tight loops
Thresholds
Section titled “Thresholds”The thresholds below are intentionally broad. They describe the approximate cost of one call, not the total time spent running the task:
- Best: < 1 µs
- Fast: < 2 µs
- Good: < 4 µs
- Fair: < 9 µs
- Slow: > 9 µs
Benchmark context
Section titled “Benchmark context”The figures on this page were measured on an Apple M3 Ultra running Node 24.12.0 on arm64-darwin, at roughly 3.86 GHz. Treat them as useful comparisons, not promises: your runtime, CPU, payload shape, and task itself will change the result.
How payloads move
Section titled “How payloads move”Most performance differences come down to how much data has to cross the worker boundary and whether that data is copied.
- Small values. Numbers, booleans, short strings, and similar values fit in the call header, so encoding and decoding are very cheap.
- Small binary payloads. Typed arrays and other small values use the transport’s preallocated space and usually stay fast.
- Larger payloads. The transport may need to allocate more space and copy the data. This is still efficient, but the cost becomes visible in a hot loop.
- Shared memory.
SharedArrayBufferandProcessSharedBufferavoid copying the bytes. Knitting passes a handle instead, so the cost of sending a 1 KiB buffer and a 64 MiB buffer is roughly the same.BufferReference(knitting/unsafe) is the thread-only move variant: it detaches the source and hands the bytes to the worker. For large top-level thread-worker returns, ordinaryUint8ArrayandArrayBuffervalues use the safe ownership path automatically at 256 KiB and above.
Typical cost by value type
Section titled “Typical cost by value type”These ratings describe one value passed in a single call. They are useful for choosing a representation, but measure your real workload before tuning around them. Types that carry a variable amount of data appear at more than one size, because their cost follows the bytes they hold instead of settling on a single rating.
| Value | Typical cost |
|---|---|
Primitives: boolean, undefined, null | Best |
Numbers: number | Best |
Time/IDs: Date | Best |
Strings: small string | Best |
Symbols: Symbol.for | Fast |
BigInt: bigint up to 64 bits | Best |
BigInt: bigint of a few hundred bits | Good |
BigInt: bigint of thousands of bits | Slow |
| Binary: typed arrays up to ~1 KiB | Best |
| Binary: typed arrays of tens of KiB | Fair |
| Binary: typed arrays of 1 MiB or more | Slow |
Views: DataView | Good |
| Structured: JSON object | Good |
| Structured: JSON array | Good |
Errors: Error | Slow |
Tuning the pool
Section titled “Tuning the pool”Thread count
Section titled “Thread count”Thread count depends on what the host thread must still do. For a mixed HTTP
service, start with one worker: the host still accepts connections, routes,
encodes requests, and sends responses. Add workers only when CPU-heavy calls
queue and lower tail latency is worth the extra coordination. For independent
batch compute, os.availableParallelism() - 1 is a reasonable first trial.
Every worker also consumes memory for payload buffers, shared locks, and cancellation state. Adding threads beyond the available cores usually stops helping. See Multi-threading for a server-focused selection process and timer configuration.
Native work stealing
Section titled “Native work stealing”Compatible multi-worker pools use native work stealing by default. The host publishes calls to one shared submit region, and workers claim available tasks as they become free. Responses still use private return lanes, so the task API and promise behavior do not change.
This helps most when many CPU-bound calls compete for workers and task durations
are uneven. It is unlikely to help much with tasks that mostly wait on
databases, networks, or other external I/O. Set host.steal: false for a
private-lane baseline, or tune host.stealRegionLanes for the workload:
- Wider regions reduce arbitration overhead for many cheap, similarly sized calls.
- Narrower regions expose more independent work for expensive or uneven calls.
1is a useful starting point when one task can take much longer than another.
Why the default can look slower
Section titled “Why the default can look slower”A benchmark that calls one function with one cost, over and over, is the worst case for stealing. Every call takes the same time, so private lanes can hand work out in strict rotation and never guess wrong, while the shared region pays for arbitration it gains nothing from. On that shape of workload we have measured a noticeably worse p99 with stealing on. If your workload really is that shape, numerical kernels and other evenly divided parallel math, turning stealing off is a fair thing to try.
Most services are not that shape. As soon as calls differ in cost, whether it is one function with an uneven input or two functions with very different prices, private lanes start binding a cheap call to a worker that is midway through an expensive one. The cheap call waits for work it has nothing to do with. Stealing has no such thing as the wrong worker: whichever one frees up first takes the next call. In a mix of a quarter-millisecond function and a twenty-millisecond function, we have seen the median call settle several times faster with stealing on, and the gap widens as load rises.
Two things follow from that, and both matter more than any single percentile:
- It holds up under load. The advantage is small when workers are mostly idle, since there is little queued work to move around. It grows as the pool fills, and near saturation private lanes are the first to build a backlog they cannot clear.
- It usually costs less CPU. A worker with nothing to claim parks instead of spinning, so the same work tends to finish for less total CPU time. The saving is largest when the pool is only partly busy, which is where most services sit.
The honest summary is that stealing trades a little uniform-workload speed for predictability everywhere else. Since the pools that benefit are the common ones, it is on by default, and a workload uniform enough to prefer private lanes is usually obvious from the code.
The host doorbell is a separate completion optimization. Node.js and Bun
thread pools use Atomics.waitAsync; Deno can use a thread-safe FFI callback,
and process workers use a process-local completion transport. Denied or
unsupported configurations fall back to polling; compiled workers and browsers
remain on their documented fallback paths. Compare like with like when
benchmarking: keep the worker topology fixed and change one of host.steal or
host.doorbell at a time. See Work stealing for the
full option and compatibility details.
Inliner
Section titled “Inliner”The inliner runs eligible calls without sending them through a worker, skipping
transport encoding and decoding. It is most useful for very small tasks, such
as simple arithmetic. For example,
inliner: { position: "last", batchSize: 64 } can improve throughput. See the
Inliner guide.
Permissions
Section titled “Permissions”Strict permissions add a small amount of startup work for each worker, such as generating flags and resolving lock files. They do not add measurable overhead to individual calls once the workers are running.
Payload sizing
Section titled “Payload sizing”payload.payloadInitialBytes, payload.payloadMaxByteLength, and
payload.maxPayloadBytes control the transport buffer used by each worker.
Increasing the initial size avoids growth later, but uses more memory up front.
The defaults are 4 MiB initially, 64 MiB of maximum length, and an 8 MiB
cap on any single dynamic payload. For consistently small payloads such as
primitives and short strings, that is usually enough.
maxPayloadBytes is a limit rather than a slow path. A value that encodes
larger than the cap throws instead of growing the buffer, so a route that
occasionally receives a large body needs either a higher cap or a
representation that travels by reference.
Choosing threads or processes
Section titled “Choosing threads or processes”Threads and processes have different memory boundaries, so the best zero-copy option depends on which one you use:
| Runtime | Isolation | Zero-copy tools |
|---|---|---|
thread (default) | shares the host address space | SharedArrayBuffer, BufferReference (move) |
process | separate memory and permissions | ProcessSharedBuffer (OS shared memory) |
A SharedArrayBuffer or BufferReference cannot cross a process boundary. Use
ProcessSharedBuffer when the worker runs in a separate process. See
Shared memory and
Buffer reference.
Because a process worker is a child process, you can start it through another
tool using worker.processCommandPrefix. This is useful for a sandbox such as
bwrap or a container such as Docker:
worker: { runtime: "process", processCommandPrefix: ["bwrap", "--unshare-all", "--ro-bind", "/", "/"],}See Process workers for the full wrapper recipes.
Keep request bodies off the main thread
Section titled “Keep request bodies off the main thread”call.*() accepts promises for supported inputs. In an HTTP handler, you can
pass the request body’s promise straight to a task instead of awaiting it on
the request thread:
app.post("/jwt", async (c) => { const responseJson = await handlers.call.issueJwt(c.req.arrayBuffer());
return c.body(responseJson ?? "Bad request", responseJson ? 200 : 400, { "content-type": "application/json; charset=utf-8", });});Knitting awaits the body promise on the host before dispatch. Passing it directly simplifies the handler while keeping decoding and parsing in the worker:
c.req.arrayBuffer()already returns a promise, so forwarding it skips anawaitin the handler.- UTF-8 decoding and JSON parsing happen in the worker, not on the request thread.
ArrayBufferstays on the binary fast path.
Metadata and body with Envelope
Section titled “Metadata and body with Envelope”When a task needs both request metadata and the raw body, put them in an
Envelope. The header carries the metadata and the payload carries the bytes.
Use .then(...) to build the envelope from the body promise without awaiting
the body on the request thread:
import { Envelope } from "knitting";
app.post("/upload", async (c) => { const result = await handlers.call.storeUpload( c.req.arrayBuffer().then( (body) => new Envelope( { contentType: c.req.header("content-type") ?? "application/octet-stream" }, body, ), ), );
return c.json(result);});For a large binary body sent to a thread worker, use a BufferReference
instead of an ArrayBuffer. The bytes move without a copy; only the body line
changes:
import { BufferReference } from "knitting/unsafe";
new Envelope( { contentType: c.req.header("content-type") ?? "application/octet-stream" }, new BufferReference(body), // moves the body bytes without copying);This works best when the route is mainly forwarding data and the worker does the parsing, such as SSR or JWT issuance. If the main thread needs to inspect or validate the body first, await it there instead.
Large bodies with allocOrRefer()
Section titled “Large bodies with allocOrRefer()”Choosing between an ArrayBuffer and a BufferReference by hand means picking
a size threshold and keeping it honest as the workload shifts. allocOrRefer()
from knitting/shared-memory makes that call for each body instead. Small
bodies are written into the allocator’s arena, larger ones are moved into a
BufferReference, and either way the task receives one body.wire value:
import { createKnittingAllocator } from "knitting/shared-memory";
const allocator = createKnittingAllocator({ arenaByteLength: 8 * 1024 * 1024 });
app.post("/upload", async (c) => { using body = await allocator.allocOrRefer(c.req.raw, { maxByteLength: 8 * 1024 * 1024, referenceAboveBytes: 2 * 1024 * 1024, });
return c.json(await handlers.call.storeUpload(body.wire));});The worker attaches to the arena once through worker.bootstrap and then reads
every body the same way, whichever transport it arrived on. See HTTP body
helpers for the bootstrap module and
the ownership rules that come with the handle.
What this buys is a send cost that stops following the body size. A plain typed
array is encoded into the call, so a 1 MiB body is far more expensive to send
than a 1 KiB one. body.wire points at bytes both sides can already reach, so
it sends in roughly the same time at any size. The arena has bookkeeping of its
own, so for bodies of about a kilobyte a plain ArrayBuffer is still the
cheaper choice.
The crossover sizes are exported, so a handler can reference them instead of repeating the numbers:
import { HTTP_BODY_REFERENCE_THRESHOLD_BYTES, // 2 MiB: move into a BufferReference HTTP_BODY_STREAM_THRESHOLD_BYTES, // 192 KiB: stream straight into the arena} from "knitting/shared-memory";Choosing how to return data
Section titled “Choosing how to return data”Returning data has the same costs as sending it, just in the other direction. Choose the return type based on the size of the result and how the data is already represented:
- JSON object / array — serialized on the worker and parsed again on the host, so it makes two passes over the data. This is fine for small results but expensive for large ones.
SharedArrayBuffer/ProcessSharedBuffer— shared memory is usually the cheapest way to return bytes when the result can use it.ProcessSharedBufferalso works with process workers.- Large plain byte returns — return an ordinary top-level
Uint8ArrayorArrayBuffer; at 256 KiB and above the thread worker uses the safe ownership path automatically. Node with the native addon can adopt the backing store, while Deno, Bun, and older Node backends make one private copy. The result remains valid after later calls and pool shutdown. The switch is worth knowing about when a result lands near it: a return just under 256 KiB is encoded into the call, while one at or above it takes the ownership path and can arrive faster than the smaller result did. BufferReference— use it mainly for large inputs that are already in an ownedArrayBufferor typed-array view and need to move into a thread worker. For an intentionally short-lived borrowed return, usesharedBytes()withunsafe: { SharedBytes: true }; see the Buffer reference guide.
See Payloads and Buffer reference for the full type list.