Skip to content

Benchmarks

This page gives you the short version of Knitting’s benchmark results across Node, Deno, Bun, and Tokio.

The tests focus on worker communication, payload sizes, batching, and CPU-heavy work. Use them to compare workloads and transport costs. Your application’s performance will depend on its own workload and environment.

The Node, Deno, and Bun figures come from Knitting 0.1.24 on an Apple M3 Ultra, recorded in March 2026. The Tokio comparison is a separate run on different hardware, described on its own page.

  • Small calls have low overhead.
  • Batching improves throughput, especially as the number of calls grows.
  • Large binary payloads are mainly limited by copying and memory bandwidth.
  • CPU-heavy work scales across workers, but the result depends on the runtime and the number of threads.

When the charts say one option is “faster,” that only applies to the test shown: the payload, batching, runtime, and amount of work all matter.



Tokio close-up

Primitive handoffs, copy-heavy bytes, and a separate Arc reference

This close-up mixes two kinds of communication cost: primitive handoff slices and one copy-heavy Uint8Array slice from the fairer default tables. The Arc<Vec<u8>> card is separate and should be read as a Tokio shared-ownership reference, not as the default apples-to-apples byte comparison. Markers to the left are faster, markers to the right are slower.

number f64 · batch 1

Primitive handoff only: Node and Bun are faster, Tokio stays ahead of Deno, and all four remain in the 6-22 µs range.

  • Tokio 13.01 µs baseline
  • Bun + Knitting 7.35 µs 44% faster
  • Node + Knitting 6.63 µs 49% faster
  • Deno + Knitting 21.54 µs 66% slower

number f64 · batch 10

Still a handoff-cost slice: Bun, Node, and Deno stay below 20 µs, while Tokio averages 27.50 µs.

  • Tokio 27.50 µs baseline
  • Bun + Knitting 13.41 µs 51% faster
  • Node + Knitting 17.28 µs 37% faster
  • Deno + Knitting 11.97 µs 56% faster

Uint8Array 512 KiB · batch 100

This is the copy/clone-heavy byte path: Bun and Node edge Tokio on average, while Deno is a bit slower and all four land in the same 22-25 ms band.

  • Tokio 23.06 ms baseline
  • Bun + Knitting 22.33 ms 3% faster
  • Node + Knitting 22.66 ms 2% faster
  • Deno + Knitting 25.26 ms 10% slower

Arc<Vec<u8>> ref · 512 B · batch 100

Separate Tokio shared-ownership reference, not the default byte benchmark: Bun stays within 6% of Tokio, while Node and Deno are slower.

  • Tokio 79.51 µs baseline
  • Bun + Knitting 74.78 µs 6% faster
  • Node + Knitting 97.23 µs 22% slower
  • Deno + Knitting 123.11 µs 55% slower

Primitive and copied-byte rows come from the fairer default comparison. The Arc row is included as a separate small-payload shared-ownership reference.

This chart compares Knitting with worker postMessage, WebSocket, and HTTP.

In these runs, Knitting is roughly:

  • 3.3–12.5× faster than worker postMessage;
  • 4–12× faster than WebSocket;
  • 10.5–62× faster than HTTP.
Combined IPC benchmark chart across runtimes

This chart shows what happens as each test sends more messages at a time.

Knitting is generally around 3–45× faster than the worker baselines in these runs. The advantage is clearest in small and medium batches, where communication overhead makes up more of the total time.

Latency line chart across runtimes

These charts compare one value at a time with batches of 100 values.

Structured and binary values cost more than primitives, and large objects and errors cost more again. Batching improves throughput while keeping Knitting’s relative advantage in these tests.

Combined types benchmark chart for count 1 Combined types benchmark chart for count 100

These are one-way transfer results for a 1 MiB payload with a batch size of 64:

RuntimeString (GB/s)Uint8Array (GB/s)
Node1.447.50
Deno3.375.78
Bun11.8616.21

Bun is fastest in this particular 1 MiB test for both strings and binary data. At this size, memory bandwidth and runtime details matter as much as the worker coordination itself.

This test distributes a CPU-intensive prime-number workload across extra threads. In these runs, speedup reaches roughly 3.5–3.8× with four extra threads, while efficiency stays around 70–77%.

That does not mean every application will scale the same way. I/O-heavy work usually behaves very differently.

Heavy-load speedup chart across runtimes Heavy-load efficiency chart across runtimes

This page is for the broad trends. For raw tables and runtime-specific notes, see the dedicated Node, Deno, Bun, and Tokio pages.

The harness lives in the runtime repository, mimiMonads/knitting, under bench/. Clone it and run the driver from the repository root:

Terminal window
./run.sh

Results are written to results/, one Markdown file per benchmark. To emit JSON instead, for plotting or for your own scripts:

Terminal window
./run.sh --json

The Tokio comparison has its own harness in mimiMonads/knitting-vs-tokio-bench.