Skip to content

Performance

These ratings are a rough guide to the cost of one call. A Slow entry can still be 2–4× faster than postMessage, depending on the payload and the work being done.

  • Best — almost no per-call overhead
  • Fast — inexpensive for most workloads
  • Good — suitable for normal workloads
  • Fair — fine occasionally; watch hot paths
  • Slow — consider another representation in tight loops

The thresholds below are intentionally broad. They describe the approximate cost of one call, not the total time spent running the task:

  • Best: < 1 µs
  • Fast: < 2 µs
  • Good: < 4 µs
  • Fair: < 9 µs
  • Slow: > 9 µs

The figures on this page were measured on an Apple M3 Ultra running Node 24.12.0 on arm64-darwin, at roughly 3.86 GHz. Treat them as useful comparisons, not promises: your runtime, CPU, payload shape, and task itself will change the result.


Most performance differences come down to how much data has to cross the worker boundary and whether that data is copied.

  • Small values. Numbers, booleans, short strings, and similar values fit in the call header, so encoding and decoding are very cheap.
  • Small binary payloads. Typed arrays and other small values use the transport’s preallocated space and usually stay fast.
  • Larger payloads. The transport may need to allocate more space and copy the data. This is still efficient, but the cost becomes visible in a hot loop.
  • Shared memory. SharedArrayBuffer and ProcessSharedBuffer avoid copying the bytes. Knitting passes a handle instead, so the cost of sending a 1 KiB buffer and a 64 MiB buffer is roughly the same. BufferReference (knitting/unsafe) is the thread-only move variant: it detaches the source and hands the bytes to the worker. For large top-level thread-worker returns, ordinary Uint8Array and ArrayBuffer values use the safe ownership path automatically at 256 KiB and above.

These ratings describe one value passed in a single call. They are useful for choosing a representation, but measure your real workload before tuning around them. Types that carry a variable amount of data appear at more than one size, because their cost follows the bytes they hold instead of settling on a single rating.

ValueTypical cost
Primitives: boolean, undefined, nullBest
Numbers: numberBest
Time/IDs: DateBest
Strings: small stringBest
Symbols: Symbol.forFast
BigInt: bigint up to 64 bitsBest
BigInt: bigint of a few hundred bitsGood
BigInt: bigint of thousands of bitsSlow
Binary: typed arrays up to ~1 KiBBest
Binary: typed arrays of tens of KiBFair
Binary: typed arrays of 1 MiB or moreSlow
Views: DataViewGood
Structured: JSON objectGood
Structured: JSON arrayGood
Errors: ErrorSlow

Thread count depends on what the host thread must still do. For a mixed HTTP service, start with one worker: the host still accepts connections, routes, encodes requests, and sends responses. Add workers only when CPU-heavy calls queue and lower tail latency is worth the extra coordination. For independent batch compute, os.availableParallelism() - 1 is a reasonable first trial.

Every worker also consumes memory for payload buffers, shared locks, and cancellation state. Adding threads beyond the available cores usually stops helping. See Multi-threading for a server-focused selection process and timer configuration.

Compatible multi-worker pools use native work stealing by default. The host publishes calls to one shared submit region, and workers claim available tasks as they become free. Responses still use private return lanes, so the task API and promise behavior do not change.

This helps most when many CPU-bound calls compete for workers and task durations are uneven. It is unlikely to help much with tasks that mostly wait on databases, networks, or other external I/O. Set host.steal: false for a private-lane baseline, or tune host.stealRegionLanes for the workload:

  • Wider regions reduce arbitration overhead for many cheap, similarly sized calls.
  • Narrower regions expose more independent work for expensive or uneven calls. 1 is a useful starting point when one task can take much longer than another.

A benchmark that calls one function with one cost, over and over, is the worst case for stealing. Every call takes the same time, so private lanes can hand work out in strict rotation and never guess wrong, while the shared region pays for arbitration it gains nothing from. On that shape of workload we have measured a noticeably worse p99 with stealing on. If your workload really is that shape, numerical kernels and other evenly divided parallel math, turning stealing off is a fair thing to try.

Most services are not that shape. As soon as calls differ in cost, whether it is one function with an uneven input or two functions with very different prices, private lanes start binding a cheap call to a worker that is midway through an expensive one. The cheap call waits for work it has nothing to do with. Stealing has no such thing as the wrong worker: whichever one frees up first takes the next call. In a mix of a quarter-millisecond function and a twenty-millisecond function, we have seen the median call settle several times faster with stealing on, and the gap widens as load rises.

Two things follow from that, and both matter more than any single percentile:

  • It holds up under load. The advantage is small when workers are mostly idle, since there is little queued work to move around. It grows as the pool fills, and near saturation private lanes are the first to build a backlog they cannot clear.
  • It usually costs less CPU. A worker with nothing to claim parks instead of spinning, so the same work tends to finish for less total CPU time. The saving is largest when the pool is only partly busy, which is where most services sit.

The honest summary is that stealing trades a little uniform-workload speed for predictability everywhere else. Since the pools that benefit are the common ones, it is on by default, and a workload uniform enough to prefer private lanes is usually obvious from the code.

The host doorbell is a separate completion optimization. Node.js and Bun thread pools use Atomics.waitAsync; Deno can use a thread-safe FFI callback, and process workers use a process-local completion transport. Denied or unsupported configurations fall back to polling; compiled workers and browsers remain on their documented fallback paths. Compare like with like when benchmarking: keep the worker topology fixed and change one of host.steal or host.doorbell at a time. See Work stealing for the full option and compatibility details.

The inliner runs eligible calls without sending them through a worker, skipping transport encoding and decoding. It is most useful for very small tasks, such as simple arithmetic. For example, inliner: { position: "last", batchSize: 64 } can improve throughput. See the Inliner guide.

Strict permissions add a small amount of startup work for each worker, such as generating flags and resolving lock files. They do not add measurable overhead to individual calls once the workers are running.

payload.payloadInitialBytes, payload.payloadMaxByteLength, and payload.maxPayloadBytes control the transport buffer used by each worker. Increasing the initial size avoids growth later, but uses more memory up front. The defaults are 4 MiB initially, 64 MiB of maximum length, and an 8 MiB cap on any single dynamic payload. For consistently small payloads such as primitives and short strings, that is usually enough.

maxPayloadBytes is a limit rather than a slow path. A value that encodes larger than the cap throws instead of growing the buffer, so a route that occasionally receives a large body needs either a higher cap or a representation that travels by reference.


Threads and processes have different memory boundaries, so the best zero-copy option depends on which one you use:

RuntimeIsolationZero-copy tools
thread (default)shares the host address spaceSharedArrayBuffer, BufferReference (move)
processseparate memory and permissionsProcessSharedBuffer (OS shared memory)

A SharedArrayBuffer or BufferReference cannot cross a process boundary. Use ProcessSharedBuffer when the worker runs in a separate process. See Shared memory and Buffer reference.

Because a process worker is a child process, you can start it through another tool using worker.processCommandPrefix. This is useful for a sandbox such as bwrap or a container such as Docker:

worker: {
runtime: "process",
processCommandPrefix: ["bwrap", "--unshare-all", "--ro-bind", "/", "/"],
}

See Process workers for the full wrapper recipes.


call.*() accepts promises for supported inputs. In an HTTP handler, you can pass the request body’s promise straight to a task instead of awaiting it on the request thread:

app.post("/jwt", async (c) => {
const responseJson = await handlers.call.issueJwt(c.req.arrayBuffer());
return c.body(responseJson ?? "Bad request", responseJson ? 200 : 400, {
"content-type": "application/json; charset=utf-8",
});
});

Knitting awaits the body promise on the host before dispatch. Passing it directly simplifies the handler while keeping decoding and parsing in the worker:

  • c.req.arrayBuffer() already returns a promise, so forwarding it skips an await in the handler.
  • UTF-8 decoding and JSON parsing happen in the worker, not on the request thread.
  • ArrayBuffer stays on the binary fast path.

When a task needs both request metadata and the raw body, put them in an Envelope. The header carries the metadata and the payload carries the bytes. Use .then(...) to build the envelope from the body promise without awaiting the body on the request thread:

import { Envelope } from "knitting";
app.post("/upload", async (c) => {
const result = await handlers.call.storeUpload(
c.req.arrayBuffer().then(
(body) =>
new Envelope(
{ contentType: c.req.header("content-type") ?? "application/octet-stream" },
body,
),
),
);
return c.json(result);
});

For a large binary body sent to a thread worker, use a BufferReference instead of an ArrayBuffer. The bytes move without a copy; only the body line changes:

import { BufferReference } from "knitting/unsafe";
new Envelope(
{ contentType: c.req.header("content-type") ?? "application/octet-stream" },
new BufferReference(body), // moves the body bytes without copying
);

This works best when the route is mainly forwarding data and the worker does the parsing, such as SSR or JWT issuance. If the main thread needs to inspect or validate the body first, await it there instead.

Choosing between an ArrayBuffer and a BufferReference by hand means picking a size threshold and keeping it honest as the workload shifts. allocOrRefer() from knitting/shared-memory makes that call for each body instead. Small bodies are written into the allocator’s arena, larger ones are moved into a BufferReference, and either way the task receives one body.wire value:

import { createKnittingAllocator } from "knitting/shared-memory";
const allocator = createKnittingAllocator({ arenaByteLength: 8 * 1024 * 1024 });
app.post("/upload", async (c) => {
using body = await allocator.allocOrRefer(c.req.raw, {
maxByteLength: 8 * 1024 * 1024,
referenceAboveBytes: 2 * 1024 * 1024,
});
return c.json(await handlers.call.storeUpload(body.wire));
});

The worker attaches to the arena once through worker.bootstrap and then reads every body the same way, whichever transport it arrived on. See HTTP body helpers for the bootstrap module and the ownership rules that come with the handle.

What this buys is a send cost that stops following the body size. A plain typed array is encoded into the call, so a 1 MiB body is far more expensive to send than a 1 KiB one. body.wire points at bytes both sides can already reach, so it sends in roughly the same time at any size. The arena has bookkeeping of its own, so for bodies of about a kilobyte a plain ArrayBuffer is still the cheaper choice.

The crossover sizes are exported, so a handler can reference them instead of repeating the numbers:

import {
HTTP_BODY_REFERENCE_THRESHOLD_BYTES, // 2 MiB: move into a BufferReference
HTTP_BODY_STREAM_THRESHOLD_BYTES, // 192 KiB: stream straight into the arena
} from "knitting/shared-memory";

Returning data has the same costs as sending it, just in the other direction. Choose the return type based on the size of the result and how the data is already represented:

  • JSON object / array — serialized on the worker and parsed again on the host, so it makes two passes over the data. This is fine for small results but expensive for large ones.
  • SharedArrayBuffer / ProcessSharedBuffer — shared memory is usually the cheapest way to return bytes when the result can use it. ProcessSharedBuffer also works with process workers.
  • Large plain byte returns — return an ordinary top-level Uint8Array or ArrayBuffer; at 256 KiB and above the thread worker uses the safe ownership path automatically. Node with the native addon can adopt the backing store, while Deno, Bun, and older Node backends make one private copy. The result remains valid after later calls and pool shutdown. The switch is worth knowing about when a result lands near it: a return just under 256 KiB is encoded into the call, while one at or above it takes the ownership path and can arrive faster than the smaller result did.
  • BufferReference — use it mainly for large inputs that are already in an owned ArrayBuffer or typed-array view and need to move into a thread worker. For an intentionally short-lived borrowed return, use sharedBytes() with unsafe: { SharedBytes: true }; see the Buffer reference guide.

See Payloads and Buffer reference for the full type list.