Skip to content

Work stealing

Knitting has two independent host-side scheduling features:

  • Native work stealing changes how requests reach workers. Compatible multi-worker pools publish work to one shared submit region, so workers can claim tasks as they become available.
  • The host doorbell changes how the host waits for responses. Supported runtimes can arm an asynchronous waiter instead of repeatedly polling the response mailbox.

Both features reduce CPU time spent waiting. A worker with no available work parks instead of spinning, and a host with no results to collect waits on the doorbell instead of polling. This keeps idle CPU use low. Multi-threading shows what that looks like in measurements.

Neither feature changes the task API. Tasks are still exported functions or task() definitions, and calls still look like await pool.call.name(input).

import { createPool, isMain, task } from "knitting";
export const transform = task<string, string>({
f: (value) => value.toUpperCase(),
});
if (isMain) {
using pool = createPool({
threads: 4,
host: {
steal: true,
doorbell: true,
},
})({ transform });
console.log(await pool.call.transform("hello"));
}

The explicit options are useful when documenting or benchmarking a topology. For compatible multi-worker pools, native stealing is selected automatically; the task code does not need to opt in from the worker side.

With private request lanes, the host chooses a worker and publishes the call to that worker’s mailbox. A busy worker can therefore hold queued work while a different worker is idle.

With native stealing, the pool uses a shared-submit/private-return layout:

  1. The host publishes a request to one shared submit region.
  2. A worker claims a region when it can make progress on the work there.
  3. Each worker keeps a private return region for its responses.
  4. The worker that claims a task owns its response; the host’s pending registry settles the corresponding promise.

This lets a worker that finishes early claim more available work without the host predicting which worker will be free next. Completion order remains task completion order, not submission order.

Native stealing is selected automatically when the pool has multiple workers and its configuration is compatible with the shared-submit topology. It also supports process-worker pools, subject to the same shared-memory and claimant limits. A one-worker pool keeps its ordinary private lane because there is nothing to steal from. Compiled workers, inliners, explicitly private dispatcher/balancer configurations, and pools above the 31-claimant protocol limit keep their fallback topology.

An explicit balancer, private-lane dispatcher, inliner, compiled worker, or an unsupported worker count can change that compatibility decision. Use host.steal: false to force private request lanes. Use host.steal: true only when the resulting topology is supported and is what you intend to measure.

The submit region has 32 slots. Stealing divides those slots into regions, and one claiming handshake takes one whole region. stealRegionLanes is the region width:

number of regions = 32 / stealRegionLanes

Use a positive power of two. Knitting chooses the widest valid region by default, leaving a spare region alongside the worker claimants. Wider regions amortise arbitration and suit many cheap, similarly sized calls. Narrower regions expose more independent work and suit expensive or uneven calls. This setting affects throughput and load balancing, but not correctness.

host.stealClaim chooses how workers claim a region:

ValueClaim discipline
"dekker"The default. Each consumer has an intent slot and claims with a per-consumer handshake.
"cas-mask"One shared compare-and-swap mask coordinates region claims.
using pool = createPool({
threads: 4,
host: {
steal: true,
stealClaim: "cas-mask",
stealRegionLanes: 1,
},
})({ transform });

The environment variable KNITTING_STEAL_CLAIM accepts the same two values. An explicit host.stealClaim takes precedence over the environment variable. The discipline also sets the maximum useful stealRegionLanes: Dekker needs at least one spare region per live consumer, so it can constrain an explicit region width more tightly than "cas-mask".

The doorbell is a host completion mechanism. It is separate from the worker loop that waits for new requests.

When the host has drained all visible responses, a polling dispatcher schedules another notification and eventually backs off with a timer. With the doorbell, the host arms an asynchronous wait on the response mailbox’s shared signal. When a worker publishes a response, it rings that signal and the host schedules another drain. If the runtime cannot arm the wait, Knitting falls back to polling.

Polling spends host CPU whether or not a response is waiting, and it spends it on the same thread that produces the work. The doorbell removes the empty checks: nothing is scheduled until a worker has something to hand back. The less loaded the pool, the larger the share of checks that were empty, which is why the doorbell matters most on a pool that is not saturated.

The doorbell is enabled by default when host.doorbell is not false and the runtime has a completion wake path. Unsupported or denied configurations fall back to polling.

Runtime or topologyCompletion behavior
Node.js or Bun thread workersAtomics.waitAsync; Node can optionally use native uv_async_t
Deno thread workersThread-safe FFI callback when FFI permission is available; otherwise polling
Process workersProcess-local completion transport; otherwise polling
Compiled workersHost options are unsupported by the compiled-worker path
Browser web workersPolling; the browser doorbell is intentionally disabled

Set host.doorbell: false to force polling for an apples-to-apples benchmark or when the host is competing with workers for every CPU core. On Node thread workers, host.nativeDoorbell: true additionally requests the optional native uv_async_t bridge from the knitting_doorbell addon. It is off by default, requires host.doorbell, does not apply to process workers, and silently falls back when the prebuild or permission is unavailable.

Start with the defaults. Tune stealRegionLanes when task durations are highly uneven, and compare host.doorbell: false with the supported default when host CPU or completion latency matters.

Keep the topology constant while benchmarking. Do not compare a polling private-lane pool with a doorbell stealing pool and attribute the entire difference to the doorbell. These settings matter most when many calls compete for local CPU; they are unlikely to dominate a workload that mostly waits on external I/O.