Why Knitting
Your event loop can be busy while other CPU cores sit idle.
As an application grows, parsing, validation, hashing, rendering, or compression can start taking too much time on the event loop. The expensive function is doing useful work, but everything else waits behind it — including routes that were cheap all along.
Knitting gives that function another execution lane. The code stays in the application; only its execution moves to a thread or isolated process.
Keep expensive JavaScript off the event loop, not out of the application.
The problem is waiting
Section titled “The problem is waiting”CPU pressure in a JavaScript application rarely spreads evenly. A codebase may contain thousands of functions while a small number account for most of its CPU time.
On one event loop, those functions do more than slow their own callers. They
create head-of-line blocking: routing, I/O completions, and cheap requests all
wait for the same thread. A /ping handler can do almost no work and
still have high p99 latency because it arrived behind a render or hash.
When a few functions account for most of the CPU time, moving those functions to workers can be enough. The rest of the application can stay where it is.
The work doesn’t get cheaper. The waiting does.
Section titled “The work doesn’t get cheaper. The waiting does.”The Hono example makes this visible. /ssr and /jwt are CPU-heavy; /ping
returns a small response and never enters the Knitting pool.
At a fixed mixed load of 2,000 requests per second per route:
| Hono only | Hono + one worker | |
|---|---|---|
| Completed load | 6,000 rps | 5,999 rps |
| Server-process CPU | 1.08 cores | 1.10 cores |
/ping p99 | 16.93 ms | 2.31 ms |
/ssr p99 | 14.32 ms | 8.71 ms |
/jwt p99 | 25.18 ms | 9.76 ms |
An equivalent repeat of the one-worker configuration recorded 1.06 cores and a
2.18 ms /ping p99. Read the CPU result as approximately 1.1 cores for both
configurations, not as a meaningful difference between 1.08 and 1.10.
The computation did not disappear. SSR and JWT still consumed CPU; they
consumed it on another lane. Knitting kept the transport, waiting, and
coordination cost low enough that the fixed workload cost about the same CPU,
while /ping no longer queued behind the expensive work.
That is performance isolation: same application, same traffic, roughly the same CPU, but one class of work no longer controls the latency of another.
These are workload-specific 15-second runs on Bun 1.4.0, not a universal performance guarantee. The full benchmark report includes the hardware, harness, repeat variance, raw measurements, and limitations.
A function-level execution boundary
Section titled “A function-level execution boundary”The function stays in the same repository and deployment, with its types intact. You export it, hand it to a pool, and call it like an async function. What changes is where it runs and what it can touch.
import { createPool, isMain } from "knitting";
export const hello = (name: string) => "Hello " + name;
if (isMain) { using pool = createPool({})({ hello }); console.log(await pool.call.hello("World!"));}That is the boundary: one export, one pool, one call. hello still reads like
a plain function, but it no longer executes on the main thread.
The runtime handles the mechanics around that call: task discovery, routing, promises, errors, lifecycle, scheduling, and payload transport. The application chooses the few functions that need another lane rather than duplicating or extracting everything around them.
The CPU cost of the boundary is part of the product
Section titled “The CPU cost of the boundary is part of the product”Moving work to another thread only helps if communication and coordination leave enough CPU time for the application. Keeping that overhead low is a central design goal.
Useful task work should consume CPU. Waiting, polling, unnecessary copying, and oversized pools should not. That principle shapes the runtime:
- idle workers park instead of continuously spinning;
- multi-worker pools skip the single-worker spin budget and park immediately;
- supported Node.js and Bun hosts wait on a completion doorbell instead of polling empty response mailboxes;
- shared-memory mailboxes and zero-copy paths keep transport work small;
- one worker is the starting point for a mixed HTTP service, rather than one worker per visible core.
A pool intended to recover usable capacity should consume almost none when it has no work. In the Hono measurements, the current policy held a fifteen-worker pool to 0.164 idle cores; the rejected spinning policy consumed 6.368. That 39x difference shows how much CPU an idle waiting policy can consume.
The Multi-threading guide explains the shipped defaults and how to choose a worker count. The Architecture page covers the mailbox, park/wake cycle, and host doorbell.
Why not use the usual alternatives?
Section titled “Why not use the usual alternatives?”Each alternative is right for a different boundary.
Keep the work inline. This has no transport or worker machinery, and it is the right answer while the task is cheap. Once it becomes expensive, every millisecond it occupies the event loop is a millisecond unavailable to unrelated work.
Run more application processes. Cluster mode lets an application use more cores, and a load balancer can reduce some queueing. It also duplicates the whole application and gives each process the same mixed workload. Each event loop can still receive a cheap request behind an expensive one. Knitting draws the boundary between the workload classes instead: the host keeps coordination and cheap work while selected functions use worker lanes.
Hand-roll workers. This puts the boundary in the right place, but the runtime gives you a message primitive rather than a function-call system. You still own request IDs, routing, promise settlement, errors, payload encoding, lifecycle, cancellation, and scheduling. Removing that plumbing is the original reason Knitting exists.
Make the function a service. This is right when the work needs independent ownership, deployment, cross-machine scale, or a separate failure domain. It is an expensive answer when one function, owned by the same team in the same repo, merely needs another execution lane. A local worker avoids the network and deployment overhead of a service.
Knitting occupies the space between inline execution and a network service: a local execution boundary around the work that needs one.
The boundary can become stronger without changing the call
Section titled “The boundary can become stronger without changing the call”Performance isolation usually starts with a thread. Some tasks also need a trust or failure boundary. Knitting keeps the call model the same while each pool chooses how strong that boundary should be.
Trusted compute can stay on a cheap thread. A task can instead run in a separate
process with runtime permissions, a bootstrap hook, bwrap, or a container.
importTask keeps worker-only code off
the host entirely.
Knitting does not replace distributed systems. If you need cross-machine scale, independent release schedules, team ownership, or a separate network failure domain, use a service. Knitting is for the point before that: the code still belongs in the application, but the work no longer belongs on its event loop.
What Knitting does not promise
Section titled “What Knitting does not promise”- It does not make CPU-heavy work consume no CPU; it changes where that CPU is spent and keeps the boundary overhead small.
- It does not make databases, networks, or filesystems faster. Keep naturally asynchronous I/O on the event loop.
- It does not guarantee that more workers mean more throughput. A single event loop may remain the producer, so measure worker counts against the workload.
- It does not turn one benchmark into a universal cost claim. Runtime, hardware, traffic mix, payload size, and latency target all affect the result.
Knitting gives expensive functions their own execution lane while keeping communication and coordination overhead low.
Read next
Section titled “Read next”- Quick Start — the ten-line version of the execution boundary.
- Multi-threading — choose a worker count without starving the host.
- Architecture — shared-memory transport, scheduling, and park/wake behavior.
- Process workers — stronger isolation through processes, sandboxes, and containers.
- Hono server example — the mixed workload behind the latency result.