Skip to content

Parallel Monte Carlo pi

This example estimates pi by sampling random points in a square. Points inside the unit circle count as hits:

π≈4⋅points inside the circletotal points\pi \approx 4 \cdot \frac{\text{points inside the circle}}{\text{total points}}

Every point is independent, so the host can split the work into chunks and let the workers run them in parallel. Each task returns one small number — its hit count — and the host adds those counts together.

  1. The host creates a pool and divides the sample count into chunks.
  2. Each worker receives one tuple: [seed, samples].
  3. The worker runs the tight numeric loop and returns its hit count.
  4. The host reduces all counts into one estimate of pi.

The example uses a deterministic seed, so the estimate is reproducible. Change --samples to trade runtime for accuracy; Monte Carlo error falls roughly as 1 / sqrt(samples).

bun.sh
bun src/pi.ts --threads 4 --samples 50000000 --chunk 1000000

Expected output:

threads: 4
samples: 50,000,000
chunks: 50
pi: 3.14140728
error: 1.85e-4
elapsed: 900 ms

The estimate is stable for the fixed seed; elapsed time depends on your runtime, CPU, and worker count.

  • The task does one thing: simulate a chunk of points.
  • The host owns orchestration and reporting.
  • Results stay tiny instead of transferring every generated point.
  • The same pattern works for numerical integration, sampling, and simulations.

This is the map-reduce pattern in its simplest form: map independent chunks to workers, then reduce their summaries on the host.

pi.ts
import { createPool, isMain } from "knitting";
import { piChunk } from "./montecarlo_pi.ts";
type Options = {
threads: number;
samples: number;
chunk: number;
};
function positiveIntArg(name: string, fallback: number): number {
const index = process.argv.indexOf(`--${name}`);
const value = index === -1 ? undefined : Number(process.argv[index + 1]);
return Number.isInteger(value) && value > 0 ? value : fallback;
}
function readOptions(): Options {
return {
threads: positiveIntArg("threads", 4),
samples: positiveIntArg("samples", 50_000_000),
chunk: positiveIntArg("chunk", 1_000_000),
};
}
async function main() {
const options = readOptions();
const jobCount = Math.ceil(options.samples / options.chunk);
const seed = 0x1234_5678;
using pool = createPool({ threads: options.threads })({ piChunk });
const started = performance.now();
const jobs: Promise<number>[] = [];
for (let job = 0; job < jobCount; job++) {
const samples = Math.min(
options.chunk,
options.samples - job * options.chunk,
);
jobs.push(pool.call.piChunk([seed + job, samples]));
}
const inside = (await Promise.all(jobs)).reduce(
(total, count) => total + count,
0,
);
const elapsed = performance.now() - started;
const pi = (4 * inside) / options.samples;
const error = Math.abs(Math.PI - pi);
console.log(`threads: ${options.threads}`);
console.log(`samples: ${options.samples.toLocaleString()}`);
console.log(`chunks: ${jobCount.toLocaleString()}`);
console.log(`pi: ${pi.toFixed(8)}`);
console.log(`error: ${error.toExponential(2)}`);
console.log(`elapsed: ${elapsed.toFixed(0)} ms`);
}
if (isMain) {
await main();
}
  1. Lower --samples to 1_000_000 for a quick run.
  2. Increase --samples to see the estimate settle closer to pi.
  3. Compare --chunk 100000 with --chunk 5000000 and watch the scheduling trade-off.
  4. Replace the circle test with another independent simulation and keep the same pool shape.