Skip to content

arXiv corpus

This example downloads the LaTeX source of sixteen well-known arXiv papers and turns each one into a record ready for indexing or embedding: the prose with markup removed, plus the section tree, citation keys, and equation, figure and table counts.

Unlike the other examples on this site, the input here is not synthetic. It is sixteen real submissions, with everything real submissions contain — macros authors invented for themselves, files that \input other files, commented-out paragraphs, and one paper whose “source” turns out to be a PDF with a two-line LaTeX wrapper around it.

  1. The host makes sure the sixteen archives are cached, then sends each worker a file path.
  2. The worker reads its archive, gunzips it, and pulls the .tex entries out of the tar.
  3. It finds the entry point, splices in every \input, and expands the paper’s own macros.
  4. It strips comments, math, tabulars and bibliographies, and unwraps the prose macros.
  5. It returns the extracted text, the section tree, and the counters.

Five files:

  • tex_scan.ts — reading TeX: balanced groups, comments, \input, macro expansion
  • tex_parse.ts — turning a resolved document into prose and structure, and the task
  • arxiv_corpus.ts — fetching, caching, gunzip and tar, and the task worth copying
  • run_latex_papers.ts — index the corpus and print what came out
  • bench_latex_papers.ts — three placements of the same job, with mitata
bun.sh
bun src/run_latex_papers.ts

Expected output, after the first run has filled the cache:

Indexed 16 arXiv submissions on 4 workers in 162 ms
id words sec eq fig tab cites title
1910.10683 22,775 59 0 6 16 277 Exploring the Limits of Transfer Learning wi
2001.08361 10,086 44 47 24 6 76 Scaling Laws for Neural Language Models
1512.03385 7,191 14 2 7 11 150 Deep Residual Learning for Image Recognition
2010.11929 7,049 32 3 12 9 100 An Image is Worth 16x16 Words: Transformers
1810.04805 7,011 29 1 5 8 91 BERT: Pre-training of Deep Bidirectional Tra
1409.1556 6,501 22 0 0 12 87 Very Deep Convolutional Networks for Large-S
1607.06450 5,630 26 10 5 3 55 Layer Normalization
1502.03167 5,553 16 17 4 0 35 Batch Normalization: Accelerating Deep Netwo
1503.02531 5,037 17 6 0 5 11 Distilling the Knowledge in a Neural Network
1312.6114 4,901 20 30 12 0 23 Auto-Encoding Variational Bayes
1301.3781 4,727 19 5 1 8 63 Efficient Estimation of Word Representations
1706.03762 4,009 23 5 5 4 58 Attention Is All You Need
1406.2661 3,013 10 8 5 2 41 Generative Adversarial Nets
1804.02767 2,854 11 1 5 3 19 YOLOv3: An Incremental Improvement
1505.04597 2,508 7 2 4 2 23 U-Net: Convolutional Networks for Biomedical
1611.03530 0 0 0 0 0 0 (PDF-only submission)
LaTeX in : 1,089,931 bytes
prose out : 626,158 bytes
words : 98,845
macros expanded: 3,174
citations : 1,109 (551 unique keys)
PDF-only : 1 with no body to index

All three runtimes produce the same 98,845 words.

The interesting decision is where to cut the job

Section titled “The interesting decision is where to cut the job”

Indexing a paper is two steps that look very different. Unpacking the archive is I/O and decompression; parsing is string work. The obvious move is to offload the parsing, because that is the part that looks like computation.

In this benchmark, moving both steps to workers was faster. The following comparison measures all three options.

bun.sh
bun src/bench_latex_papers.ts

Expected output (4 cores / 8 threads, Bun 1.4, 16 papers per iteration):

word parity check: host=98,845 split=98,845 worker=98,845 OK match
LaTeX corpus benchmark (mitata)
workload: gunzip + untar + inline inputs + expand macros + extract prose
papers per iteration: 16
threads: 4
benchmark avg (min … max)
host: unpack + parse 102.95 ms/iter (95.50 ms … 110.16 ms)
worker: parse only, host unpacks 87.58 ms/iter (80.56 ms … 96.66 ms)
worker: unpack + parse 61.66 ms/iter (51.08 ms … 81.98 ms)
summary
worker: unpack + parse
1.42x faster than worker: parse only, host unpacks
1.67x faster than host: unpack + parse

The middle row is the one to look at. Offloading the parse — the part that looks like the real work — bought 1.18x, because the decompression stayed on the thread you were trying to free, and gunzipSync blocks it completely while it runs. Sending a path instead of a payload moves the whole job across, and the gunzip parallelizes along with everything else: 1.67x from the same pool, on the same corpus.

It is also the smaller piece of code. The host never opens an archive, so loadArchive is called in exactly one place — inside the task.

// The shape worth copying: the job crosses the boundary as a path.
export const indexArchive = task<string, PaperRecord>({
f: async (path: string) => parsePaperHost(await loadArchive(path)),
});

Sending a path lets the worker handle decompression as well as parsing. Payload transfer was a relatively small cost in this workload, as two additional measurements show.

  • Sending 64 papers’ worth of sources — 3.9 MiB of LaTeX — through a task that does nothing but count the files costs about 11 ms.
  • Returning every paper’s extracted prose (626 KB) instead of returning counters only measured the same, within noise, in both directions across thread counts.

So the site’s usual advice — return compact summaries, not large payloads — is a rule about ratios, not about sizes. When each item costs milliseconds of CPU, a 40 KB result may add no measurable overhead. When each item costs microseconds, the same result dominates. Measure the ratio for your workload rather than assuming the payload is the problem.

The first version of this example used DecompressionStream, which is the web-standard way to gunzip and works on all three runtimes. In these measurements, adding workers did not make it faster.

Gunzipping the same sixteen archives inside the pool, changing only the thread count:

threadsDecompressionStreamnode:zlib gunzipSync
143.9 ms60.8 ms
239.3 ms58.2 ms
437.1 ms32.2 ms
839.2 ms27.0 ms

DecompressionStream is flat. gunzipSync starts out slower and then scales, because it runs on the worker that called it instead of handing the work to something shared.

Most of the parser exists because of things real papers do:

  • A submission can contain no LaTeX at all. One of these sixteen is a PDF with a stub .tex that \includepdfs it. It parses to zero words, and the record says pdfOnly rather than quietly reporting a paper with no content. Adam is another, which is why it is not in the list.
  • A tarball can hold several \begin{document} files. YOLOv3 ships the paper, a rebuttal, supplementary material, and a style demo. Picking the longest file gets you the supplementary material; the paper’s own entry point is 3 KB that \inputs everything else. The parser chooses the file that produces the largest document after resolving its inputs.
  • Every author invents macros. One paper in this corpus expands 1,844 of them. Skip expansion and your index fills up with tokens nobody wrote.
  • % starts a comment, \% does not, and papers are full of commented-out drafts.

The extracted text was checked against pdftotext output from arXiv’s own PDFs, minus the bibliography: the parser lands between 0.84x and 1.17x of the reference word count across spot-checked papers. It is an indexing-grade extraction, not a typesetter.

Do not expect this to scale past a few workers

Section titled “Do not expect this to scale past a few workers”

Sixteen papers is a small corpus, and it shows. One worker performs about the same as the host. Adding workers helps clearly up to about four, and past that two sweeps on the same machine disagreed with each other — one kept improving to eight, the other got worse at six. That spread is larger than the effect, so there is no per-thread table here worth printing.

What is stable is the shape, and the reason for it. The sequential cost of the corpus is about 119 ms, and the most expensive single paper is 21 ms, so a perfect four-way split would finish in 30 ms. The pool lands closer to 55-65 ms. The corpus simply runs out: sixteen items across four workers is four items each, and the round ends when the last worker finishes.

Two practical lessons follow:

  • Size the corpus to the pool, not the pool to the machine. If you have sixteen items, four workers were enough in these measurements. Try processing more items per round before adding workers.
  • The largest item sets a floor you cannot cross. Here that floor is 21 ms and it is not yet binding. If your corpus has one item that is 10x the median, it will be, and no pool size fixes it — you would have to split that item.
arxiv_corpus.ts
import { mkdir, readFile, writeFile } from "node:fs/promises";
import { gunzipSync } from "node:zlib";
import { fileURLToPath } from "node:url";
import { task } from "knitting";
import { type PaperRecord, parsePaperHost, type PaperSource } from "./tex_parse.ts";
// Sixteen well-known papers, chosen for the spread: the smallest is 22 KB of
// LaTeX and the largest is 226 KB. Two of them are PDF-only submissions, which
// is a case any real corpus job has to survive.
export const PAPER_IDS = [
"1706.03762", // Attention Is All You Need
"1512.03385", // Deep Residual Learning
"1409.1556", // VGG
"1810.04805", // BERT
"1406.2661", // Generative Adversarial Nets
"1505.04597", // U-Net
"1312.6114", // Auto-Encoding Variational Bayes
"1502.03167", // Batch Normalization
"1301.3781", // word2vec
"1607.06450", // Layer Normalization
"2010.11929", // An Image is Worth 16x16 Words
"1503.02531", // Distilling the Knowledge in a Neural Network
"2001.08361", // Scaling Laws for Neural Language Models
"1910.10683", // T5
"1611.03530", // Rethinking Generalization
"1804.02767", // YOLOv3
];
const CACHE_DIR = fileURLToPath(new URL("./papers/", import.meta.url));
// arXiv asks that you not hammer the e-print endpoint. This downloads each
// paper once, one at a time, and every later run reads the cache. If you want
// thousands of papers, use arXiv's bulk access rather than a loop like this.
const REQUEST_SPACING_MS = 3_000;
export function archivePath(id: string): string {
return `${CACHE_DIR}${id}.tar.gz`;
}
/** Download whatever is missing from the cache. Silent when there is nothing to do. */
export async function ensureCorpus(ids: string[] = PAPER_IDS): Promise<string[]> {
await mkdir(CACHE_DIR, { recursive: true });
let fetched = 0;
for (const id of ids) {
const path = archivePath(id);
try {
await readFile(path);
continue;
} catch {
// not cached yet
}
if (fetched === 0) console.log(`Fetching sources from arXiv into ${CACHE_DIR}`);
else await new Promise((resolve) => setTimeout(resolve, REQUEST_SPACING_MS));
const response = await fetch(`https://arxiv.org/e-print/${id}`, {
headers: { "user-agent": "knitting-docs-example/1.0" },
});
if (!response.ok) throw new Error(`arXiv ${id}: HTTP ${response.status}`);
const bytes = new Uint8Array(await response.arrayBuffer());
await writeFile(path, bytes);
fetched++;
console.log(` ${id} ${bytes.byteLength.toLocaleString()} bytes`);
}
if (fetched > 0) console.log(`Cached ${fetched} new archives.\n`);
return ids.map(archivePath);
}
/**
* Read the .tex entries out of a tar archive.
*
* tar is 512-byte header blocks followed by padded contents, which is little
* enough format to hand-roll rather than take a dependency on for one example.
*/
function untarTexFiles(archive: Uint8Array): Record<string, string> {
const decoder = new TextDecoder();
const files: Record<string, string> = {};
let offset = 0;
let longName: string | null = null;
while (offset + 512 <= archive.length) {
const header = archive.subarray(offset, offset + 512);
if (header.every((byte) => byte === 0)) break;
const field = (start: number, length: number) =>
decoder.decode(header.subarray(start, start + length))
.replace(/\0.*$/, "").trim();
const prefix = field(345, 155);
const rawName = field(0, 100);
const name = longName ?? (prefix ? `${prefix}/${rawName}` : rawName);
const size = Number.parseInt(field(124, 12), 8) || 0;
const type = String.fromCharCode(header[156]!);
const body = archive.subarray(offset + 512, offset + 512 + size);
offset += 512 + Math.ceil(size / 512) * 512;
if (type === "L") {
// GNU long name: this entry's content is the next header's real name.
longName = decoder.decode(body).replace(/\0.*$/, "");
continue;
}
longName = null;
if (type !== "0" && type !== "\0") continue;
if (!/\.(tex|ltx)$/i.test(name)) continue;
files[name] = decoder.decode(body);
}
return files;
}
/** Read one archive off disk and unpack the .tex files in it. */
export async function loadArchive(path: string): Promise<PaperSource> {
const id = path.split("/").pop()!.replace(/\.tar\.gz$/, "");
const archive = await readFile(path);
// `node:zlib`, not `DecompressionStream`. Both work on all three runtimes and
// the streaming one looks more idiomatic, but it does not get faster when you
// add workers -- measured flat from 1 thread to 8. `gunzipSync` runs inside
// the worker that called it, so it scales with the pool like everything else.
return { id, files: untarTexFiles(gunzipSync(archive)) };
}
/**
* Read, unpack and parse one archive, start to finish, inside the worker.
*
* This is the shape worth copying. The job crosses the boundary as a path --
* a few dozen bytes -- rather than as the megabyte of LaTeX inside the archive,
* and the gunzip goes parallel along with the parsing instead of staying on the
* host as a serial prelude to it.
*/
export const indexArchive = task<string, PaperRecord>({
f: async (path: string) => parsePaperHost(await loadArchive(path)),
});

This pattern applies to document ingestion: files on disk or in a bucket that need decompression, parsing, and text extraction before further processing. Search indexes, RAG pipelines and dataset builds all look like this, and the per-item cost is high enough that doing it on the request thread is not an option.

The pattern to take away is the task signature. indexArchive takes a path and returns a record, which means the pool owns the whole pipeline for one item — reading included. That is both faster than splitting the pipeline across the boundary and simpler to write.