Skip to content

PDF generation

This example lays out A4 invoices with pdf-lib and returns the finished PDF bytes from a worker. Generating a batch of documents takes CPU time that can delay other requests if the layout work runs on the main thread.

pdf-lib is pure JavaScript with no native dependency, so this runs unchanged on Node, Deno and Bun.

  1. The host builds a synthetic billing run — 24 to 83 line items each, so some invoices spill onto a second page.
  2. Invoices go to the worker in batches. Each worker lays out every invoice in its batch and serializes it.
  3. The worker packs the finished documents into one buffer and returns that.
  4. The host unpacks the buffer back into individual PDFs.

Three files:

  • render_invoice.ts — the layout, the packing helpers, and the tasks
  • invoice_fixtures.ts — the synthetic billing run
  • bench_invoice_pdf.ts — host vs workers, with mitata

Only a top-level Uint8Array comes back as bytes

Section titled “Only a top-level Uint8Array comes back as bytes”

This is the constraint the example is built around, and it affects how you design tasks that return binary data.

A Uint8Array returned directly from a task arrives as a Uint8Array. The same array nested inside an object or an array does not — it arrives as a plain object with numeric keys. There is no error and no warning; the call resolves, and you get something that is no longer a typed array and may be much larger than what you sent.

// Fine: arrives as a Uint8Array.
task<Invoice, Uint8Array>({ f: renderInvoiceHost });
// Not fine: each element arrives as { "0": 37, "1": 80, ... }.
task<Invoice[], Uint8Array[]>({ f: renderBatch });

This example returns each batch as one contiguous buffer. packDocuments writes a small header — a u32 count, then a u32 length per document — and concatenates the rest; unpackDocuments reverses it with subarray, so the individual PDFs are views into the returned buffer rather than copies.

pdf-lib stamps the current time into every document, so rendering the same invoice twice produces different bytes and any parity check between host and worker fails for the wrong reason. The task pins the creation and modification dates, which makes output byte-identical run to run — worth doing in production too, if you ever want to cache or deduplicate generated documents.

using pool = createPool({ threads: 4 })({ renderInvoiceBatch });
const packed = await pool.call.renderInvoiceBatch(invoices.slice(0, 12));
const documents = unpackDocuments(packed);
await writeFile("invoice-0.pdf", documents[0]);
bun.sh
bun src/bench_invoice_pdf.ts
pdf parity check: 200 documents, host=1,114,063 bytes worker=1,114,063 bytes OK match
PDF invoice rendering benchmark (mitata)
workload: lay out an A4 invoice and serialize the PDF
invoices per iteration: 200
average document: 5.4 KB
threads: 4 | batch: 12
benchmark avg (min … max)
host (200 invoices) 874.29 ms/iter (843.70 ms … 928.11 ms)
knitting (4 threads, 200 invoices) 542.27 ms/iter (479.83 ms … 662.44 ms)
summary
knitting (4 threads, 200 invoices)
1.61x faster than host (200 invoices)

Because the documents are deterministic, the parity check compares real bytes: the same 200 invoices, the same 1,114,063 bytes, whichever side rendered them.

1.6x from four workers on four cores is a modest return, and raising threads does not fix it. Replacing this workload with plain arithmetic in the same pool, with the same batching, reached 4.8x on the same laptop. This suggests that PDF generation is limiting scaling here.

Most of the time here goes into many small drawText and drawLine calls, each allocating as it goes; document creation and font embedding are close to free by comparison. Workloads built out of a lot of small allocations tend to flatten out well before they run out of cores, and the prompt token budgeting example flattens out for a related reason.

Throughput and request latency tell different parts of the story:

  • Throughput: the pool completed the batch about 1.6x faster.
  • Request latency: moving layout work off the host frees the request thread. Laying out 200 invoices takes the better part of a second, and on the host that is a second in which nothing else is served. Moving it to a pool takes it off the request thread entirely, and that is true whether the multiplier is 1.6x or 4x.
render_invoice.ts
import { task } from "knitting";
import { PDFDocument, rgb, StandardFonts } from "pdf-lib";
export type InvoiceLine = {
description: string;
quantity: number;
unitCents: number;
};
export type Invoice = {
number: string;
customer: string;
issuedAt: string;
lines: InvoiceLine[];
};
const PAGE = { width: 595.28, height: 841.89 } as const; // A4, in points
const MARGIN = 48;
const ROW_HEIGHT = 18;
const EPOCH = new Date("2026-01-01T00:00:00Z");
const money = (cents: number) =>
(cents / 100).toLocaleString("en-US", { minimumFractionDigits: 2 });
/**
* Build a one-or-more page A4 invoice and return the PDF bytes. This is the
* whole workload: laying out rows, embedding fonts, and deflating the result.
*/
export async function renderInvoiceHost(invoice: Invoice): Promise<Uint8Array> {
const doc = await PDFDocument.create();
// pdf-lib stamps the current time into every document. Pin it, or the same
// invoice produces different bytes on every run and parity checks are useless.
doc.setCreationDate(EPOCH);
doc.setModificationDate(EPOCH);
doc.setProducer("knitting-example");
const font = await doc.embedFont(StandardFonts.Helvetica);
const bold = await doc.embedFont(StandardFonts.HelveticaBold);
let page = doc.addPage([PAGE.width, PAGE.height]);
let y = PAGE.height - MARGIN - 14;
page.drawText("INVOICE", { x: MARGIN, y, size: 22, font: bold });
y -= 26;
page.drawText(`${invoice.number} ${invoice.customer}`, {
x: MARGIN,
y,
size: 10,
font,
});
page.drawText(invoice.issuedAt, {
x: PAGE.width - MARGIN - 60,
y,
size: 10,
font,
});
y -= 28;
let totalCents = 0;
for (const line of invoice.lines) {
// Overflow onto a new page rather than drawing off the bottom edge.
if (y < MARGIN + ROW_HEIGHT * 2) {
page = doc.addPage([PAGE.width, PAGE.height]);
y = PAGE.height - MARGIN - 14;
}
const cents = line.quantity * line.unitCents;
totalCents += cents;
page.drawText(line.description, { x: MARGIN, y, size: 10, font });
page.drawText(String(line.quantity), { x: 400, y, size: 10, font });
page.drawText(money(cents), { x: 470, y, size: 10, font });
page.drawLine({
start: { x: MARGIN, y: y - 4 },
end: { x: PAGE.width - MARGIN, y: y - 4 },
thickness: 0.3,
color: rgb(0.82, 0.82, 0.82),
});
y -= ROW_HEIGHT;
}
page.drawText(`Total ${money(totalCents)}`, {
x: 400,
y: y - 10,
size: 12,
font: bold,
});
return doc.save();
}
/**
* Pack several PDFs into one buffer: a little-endian u32 count, then a u32
* length per document, then the documents back to back.
*
* This exists because only a *top-level* Uint8Array survives the return trip as
* binary. A Uint8Array nested inside an array or an object arrives as a plain
* object with numeric keys -- no error, just a silently useless and much larger
* result. So a batch has to come back as one contiguous buffer.
*/
export function packDocuments(documents: Uint8Array[]): Uint8Array {
const headerBytes = 4 + documents.length * 4;
let total = headerBytes;
for (const doc of documents) total += doc.byteLength;
const packed = new Uint8Array(total);
const header = new DataView(packed.buffer);
header.setUint32(0, documents.length, true);
let offset = headerBytes;
for (let i = 0; i < documents.length; i++) {
header.setUint32(4 + i * 4, documents[i]!.byteLength, true);
packed.set(documents[i]!, offset);
offset += documents[i]!.byteLength;
}
return packed;
}
export function unpackDocuments(packed: Uint8Array): Uint8Array[] {
const header = new DataView(packed.buffer, packed.byteOffset);
const count = header.getUint32(0, true);
const documents = new Array<Uint8Array>(count);
let offset = 4 + count * 4;
for (let i = 0; i < count; i++) {
const length = header.getUint32(4 + i * 4, true);
documents[i] = packed.subarray(offset, offset + length);
offset += length;
}
return documents;
}
/**
* Render a whole batch on the worker. One call per invoice spends more time on
* dispatch than on rendering, and one packed return keeps the bytes contiguous.
*/
export async function renderInvoiceBatchHost(
invoices: Invoice[],
): Promise<Uint8Array> {
const documents = new Array<Uint8Array>(invoices.length);
for (let i = 0; i < invoices.length; i++) {
documents[i] = await renderInvoiceHost(invoices[i]!);
}
return packDocuments(documents);
}
export const renderInvoice = task<Invoice, Uint8Array>({
f: renderInvoiceHost,
});
export const renderInvoiceBatch = task<Invoice[], Uint8Array>({
f: renderInvoiceBatchHost,
});

Invoices, statements, shipping labels, exam papers, contracts — anything a product generates in bulk on a schedule, or on demand while a user waits. The work is CPU-bound and produces bytes you have to hand back, making it a useful workload for a worker pool. Pay attention to the return type: one contiguous buffer per call, not a structure with typed arrays buried inside it.