Tokio
This page compares Tokio and Knitting on Bun, Node.js, and Deno using the same batch-oriented echo benchmark.
JavaScript runtimes are quicker for some tiny values, while Tokio becomes much more competitive as the payload gets larger. The charts below show where that change happens.
Benchmark source: mimiMonads/knitting-vs-tokio-bench.
What this benchmark measures
Section titled “What this benchmark measures”Each test sends a value to a worker, echoes it back, and measures the full batch round trip. The payloads are:
f64values1 MiBstrings1 MiBUint8ArrayvaluesUint8Arrayvalues from8 Bto1 MiB- a separate
Arc<Vec<u8>>sweep for small shared-byte payloads
The results are from one run of Knitting 0.1.36 on Ubuntu 23.10, x86_64, with an
AMD Ryzen 7 4700U, recorded in March 2026. Every runtime uses the same batch sizes (1, 10, and 100), warmup, measured iterations, and percentile reporting.
These numbers describe this benchmark and this machine, not a universal ranking of runtimes.
Small payloads: f64
Section titled “Small payloads: f64”For small scalar values, the JavaScript runtimes are faster on average in this run:
- At
n=1, Node.js is lowest at6.63 µs, followed by Bun, Tokio, and Deno. - At
n=10, Deno is lowest at11.97 µs; Bun and Node.js follow, with Tokio at27.50 µs. - At
n=100, Node.js and Deno remain ahead of Tokio, while Bun sits between them.
Tokio does have the best p99 at n=1. Bun has the best p99 at n=10 and n=100.
The chart uses a logarithmic y-axis. Lower is better; the exact values are in the raw report at the end of this page.
Large payloads: 1 MiB
Section titled “Large payloads: 1 MiB”The picture changes once each message is 1 MiB:
- Tokio is fastest for the large-string test at
n=1andn=100. Bun is just ahead atn=10. - Tokio leads the
Uint8Arraytest at every batch size:272.81 µs,4.64 ms, and37.83 ms. - Node.js and Deno fall further behind as payload materialization becomes a larger part of the round trip.
These timings include the trip to the worker and the trip back, including the work needed to materialize the payload on both sides.
How the payload size changes things
Section titled “How the payload size changes things”This sweep keeps the batch at 100 and grows a binary payload from 8 B to 1 MiB.
Bun is fastest from 8 B through 512 B. Tokio takes the lead from 1 KiB through 16 KiB, then Bun and Node.js move slightly ahead through 512 KiB. At 1 MiB, Tokio is fastest again at 29.48 ms.
Both axes are logarithmic, so the chart is showing changes across several orders of magnitude. The key trend is the transition from fixed per-message overhead to payload-copying cost.
A separate shared-ownership reference
Section titled “A separate shared-ownership reference”The next chart is intentionally different. Tokio uses Arc<Vec<u8>>, so sending a value mostly increments a reference count instead of copying the bytes. This is useful as an upper-bound reference for shared ownership, but it is not the default apples-to-apples byte comparison.
At 8 B through 256 B, Bun is still faster than the Tokio Arc path. At 512 B, the two are close: 74.78 µs for Bun versus 79.51 µs for Tokio.
Read this chart as “how close can normal transport get to shared ownership?” For the fair default comparison, use the regular Uint8Array sweep above.
How fair is the comparison?
Section titled “How fair is the comparison?”The benchmark keeps the important parts aligned:
- Both sides create the work up front and wait for the whole batch to finish.
- Both use one worker thread, so both process the batch with the same worker count.
- The default string and byte paths pay for a full round trip. Tokio clones on send and clones again for the worker reply rather than transferring ownership of the original allocation.
The deliberate difference is memory management. Tokio’s String and Vec<u8> paths allocate and copy with clone(). Knitting copies into preallocated shared memory and manages those regions itself. That cost is part of what this benchmark is trying to measure, so it is left in the numbers rather than normalized away.
The Arc<Vec<u8>> test is kept separate because shared ownership does less work than the normal byte path.
Why the results change with payload size
Section titled “Why the results change with payload size”Knitting keeps transport data in reused typed-array-backed buffers. Small primitive values fit in the per-call header, while larger values spill into a preallocated shared payload buffer. Fixed worker lanes and a small allocation bookkeeping layer reduce general-purpose allocation in the hot path.
That design has a tradeoff: it adds implementation complexity and memory bookkeeping. The payoff shows up most clearly when copying and allocating a large payload would otherwise dominate the round trip.
Full raw report
Section titled “Full raw report”The generated report contains the exact average and p99 tables, ratios, machine details, and methodology notes used for the charts.
# Benchmark Summary
## Sources
- tokio: `results/tokio-1773827721825.csv`- bun: `results/knitting-bun-1773827812276.csv`- node: `results/knitting-node-1773828074699.csv`- deno: `results/knitting-deno-1773827925019.csv`
## Machine Specs
- OS: Ubuntu 23.10- Kernel: 6.5.0-44-generic- Architecture: x86_64- CPU: AMD Ryzen 7 4700U with Radeon Graphics- Topology: 8 logical CPUs, 1 socket(s), 8 core(s)/socket, 1 thread(s)/core- Memory: 15.1 GiB- Swap: 4.0 GiB
## Methodology Notes
- The main string and byte benchmarks are intended to compare the same logical round trip on both sides: send payload, receive it in the worker, echo it back, receive it again on the caller, then wait for the whole batch.- In `src/main.ts`, the `string` and `Uint8Array` paths go through knitting transport in both directions. That transport materializes a fresh payload on receive, so the round trip includes payload work on both the request side and the reply side.- To keep the Tokio baseline fair, `src/main.rs` clones `String` and `Vec<u8>` on send and also clones again on the worker reply. The reply clone is intentional. Without it, Tokio would be measuring a cheaper return-path move while the JS runtimes were still paying for fresh payload materialization on the way back.- The `Arc<Vec<u8>>` sweep is intentionally separate and is not the default apples-to-apples byte benchmark. It exists as an upper-bound shared-bytes reference for small payloads. `Arc::clone` only bumps a refcount, so it is expected to be cheaper than copying bytes.- This means the default `string` and `Uint8Array` tables should be read as the fairer comparison, while the Arc section should be read as "how close does the normal transport get to shared ownership for small values?"
## Batch Avg Latency (less is better)
```textbenchmark | batch | tokio | bun | node | deno---------------------+-------+-----------+----------+----------+---------number f64 (8 bytes) | n=1 | 13.01 us | 7.35 us | 6.63 us | 21.54 usnumber f64 (8 bytes) | n=10 | 27.50 us | 13.41 us | 17.28 us | 11.97 usnumber f64 (8 bytes) | n=100 | 89.55 us | 80.61 us | 62.33 us | 63.28 uslarge string 1 MiB | n=1 | 221.35 us | 1.19 ms | 2.85 ms | 1.38 mslarge string 1 MiB | n=10 | 6.20 ms | 6.01 ms | 10.79 ms | 10.38 mslarge string 1 MiB | n=100 | 37.93 ms | 50.16 ms | 84.66 ms | 83.90 msUint8Array 1 MiB | n=1 | 272.81 us | 1.30 ms | 2.35 ms | 1.14 msUint8Array 1 MiB | n=10 | 4.64 ms | 5.27 ms | 5.22 ms | 6.36 msUint8Array 1 MiB | n=100 | 37.83 ms | 47.76 ms | 54.95 ms | 60.04 ms```
## Batch P99 Latency (less is better)
```textbenchmark | batch | tokio | bun | node | deno---------------------+-------+-----------+----------+-----------+----------number f64 (8 bytes) | n=1 | 16.85 us | 18.70 us | 26.25 us | 160.57 usnumber f64 (8 bytes) | n=10 | 40.83 us | 36.58 us | 81.26 us | 111.89 usnumber f64 (8 bytes) | n=100 | 203.57 us | 92.37 us | 314.11 us | 263.05 uslarge string 1 MiB | n=1 | 371.10 us | 3.74 ms | 3.73 ms | 2.97 mslarge string 1 MiB | n=10 | 8.44 ms | 8.41 ms | 15.75 ms | 14.45 mslarge string 1 MiB | n=100 | 40.81 ms | 61.62 ms | 105.51 ms | 106.67 msUint8Array 1 MiB | n=1 | 400.45 us | 3.03 ms | 5.58 ms | 5.81 msUint8Array 1 MiB | n=10 | 8.41 ms | 7.80 ms | 9.52 ms | 14.37 msUint8Array 1 MiB | n=100 | 43.71 ms | 59.77 ms | 72.39 ms | 81.71 ms```
## Avg Ratio Vs Tokio
```textbenchmark | batch | bun/tokio | node/tokio | deno/tokio---------------------+-------+-----------+------------+-----------number f64 (8 bytes) | n=1 | 0.56x | 0.51x | 1.66xnumber f64 (8 bytes) | n=10 | 0.49x | 0.63x | 0.44xnumber f64 (8 bytes) | n=100 | 0.90x | 0.70x | 0.71xlarge string 1 MiB | n=1 | 5.36x | 12.86x | 6.22xlarge string 1 MiB | n=10 | 0.97x | 1.74x | 1.68xlarge string 1 MiB | n=100 | 1.32x | 2.23x | 2.21xUint8Array 1 MiB | n=1 | 4.75x | 8.62x | 4.18xUint8Array 1 MiB | n=10 | 1.13x | 1.12x | 1.37xUint8Array 1 MiB | n=100 | 1.26x | 1.45x | 1.59x```
## Uint8Array Size Sweep Avg Latency (less is better)
```textsize | tokio | bun | node | deno--------+-----------+-----------+-----------+----------8 B | 82.99 us | 62.80 us | 88.20 us | 107.01 us16 B | 81.91 us | 56.24 us | 65.37 us | 95.78 us32 B | 85.70 us | 49.48 us | 65.76 us | 85.05 us64 B | 76.98 us | 42.68 us | 66.88 us | 78.27 us128 B | 92.53 us | 53.53 us | 79.28 us | 84.39 us256 B | 99.70 us | 63.42 us | 83.89 us | 100.44 us512 B | 86.67 us | 68.55 us | 97.07 us | 118.03 us1 KiB | 101.42 us | 171.09 us | 157.61 us | 169.50 us2 KiB | 191.25 us | 194.62 us | 220.68 us | 233.39 us4 KiB | 195.56 us | 260.16 us | 324.39 us | 391.43 us8 KiB | 208.84 us | 397.05 us | 465.89 us | 539.98 us16 KiB | 279.25 us | 649.18 us | 741.47 us | 899.81 us32 KiB | 1.48 ms | 1.14 ms | 1.27 ms | 1.41 ms64 KiB | 2.71 ms | 2.38 ms | 2.49 ms | 2.89 ms128 KiB | 5.66 ms | 5.02 ms | 5.11 ms | 6.08 ms256 KiB | 11.92 ms | 10.56 ms | 10.11 ms | 11.96 ms512 KiB | 23.06 ms | 22.33 ms | 22.66 ms | 25.26 ms1 MiB | 29.48 ms | 46.97 ms | 52.53 ms | 55.77 ms```
## Arc Comparison Size Sweep Avg Latency (less is better)
Tokio uses `Arc<Vec<u8>>` here as a separate shared-bytes reference point, not the default apples-to-apples byte path.
```textsize | tokio | bun | node | deno------+----------+----------+----------+----------8 B | 80.76 us | 70.31 us | 86.19 us | 97.79 us16 B | 79.35 us | 60.94 us | 73.73 us | 77.46 us32 B | 81.48 us | 57.04 us | 70.26 us | 77.03 us64 B | 80.14 us | 54.44 us | 75.94 us | 78.81 us128 B | 79.89 us | 68.50 us | 82.51 us | 85.95 us256 B | 79.48 us | 50.59 us | 85.94 us | 100.10 us512 B | 79.51 us | 74.78 us | 97.23 us | 123.11 us```
## Arc Comparison Avg Ratio Vs Tokio
```textsize | bun/tokio | node/tokio | deno/tokio------+-----------+------------+-----------8 B | 0.87x | 1.07x | 1.21x16 B | 0.77x | 0.93x | 0.98x32 B | 0.70x | 0.86x | 0.95x64 B | 0.68x | 0.95x | 0.98x128 B | 0.86x | 1.03x | 1.08x256 B | 0.64x | 1.08x | 1.26x512 B | 0.94x | 1.22x | 1.55x```