arXiv corpus
This example downloads the LaTeX source of sixteen well-known arXiv papers and turns each one into a record ready for indexing or embedding: the prose with markup removed, plus the section tree, citation keys, and equation, figure and table counts.
Unlike the other examples on this site, the input here is not synthetic. It is sixteen real
submissions, with everything real submissions contain — macros authors invented for
themselves, files that \input other files, commented-out paragraphs, and one paper whose
“source” turns out to be a PDF with a two-line LaTeX wrapper around it.
How it works
Section titled “How it works”- The host makes sure the sixteen archives are cached, then sends each worker a file path.
- The worker reads its archive, gunzips it, and pulls the
.texentries out of the tar. - It finds the entry point, splices in every
\input, and expands the paper’s own macros. - It strips comments, math, tabulars and bibliographies, and unwraps the prose macros.
- It returns the extracted text, the section tree, and the counters.
Five files:
tex_scan.ts— reading TeX: balanced groups, comments,\input, macro expansiontex_parse.ts— turning a resolved document into prose and structure, and the taskarxiv_corpus.ts— fetching, caching, gunzip and tar, and the task worth copyingrun_latex_papers.ts— index the corpus and print what came outbench_latex_papers.ts— three placements of the same job, withmitata
bun src/run_latex_papers.tsdeno run -A src/run_latex_papers.tsnpx tsx src/run_latex_papers.tsExpected output, after the first run has filled the cache:
Indexed 16 arXiv submissions on 4 workers in 162 ms
id words sec eq fig tab cites title 1910.10683 22,775 59 0 6 16 277 Exploring the Limits of Transfer Learning wi 2001.08361 10,086 44 47 24 6 76 Scaling Laws for Neural Language Models 1512.03385 7,191 14 2 7 11 150 Deep Residual Learning for Image Recognition 2010.11929 7,049 32 3 12 9 100 An Image is Worth 16x16 Words: Transformers 1810.04805 7,011 29 1 5 8 91 BERT: Pre-training of Deep Bidirectional Tra 1409.1556 6,501 22 0 0 12 87 Very Deep Convolutional Networks for Large-S 1607.06450 5,630 26 10 5 3 55 Layer Normalization 1502.03167 5,553 16 17 4 0 35 Batch Normalization: Accelerating Deep Netwo 1503.02531 5,037 17 6 0 5 11 Distilling the Knowledge in a Neural Network 1312.6114 4,901 20 30 12 0 23 Auto-Encoding Variational Bayes 1301.3781 4,727 19 5 1 8 63 Efficient Estimation of Word Representations 1706.03762 4,009 23 5 5 4 58 Attention Is All You Need 1406.2661 3,013 10 8 5 2 41 Generative Adversarial Nets 1804.02767 2,854 11 1 5 3 19 YOLOv3: An Incremental Improvement 1505.04597 2,508 7 2 4 2 23 U-Net: Convolutional Networks for Biomedical 1611.03530 0 0 0 0 0 0 (PDF-only submission)
LaTeX in : 1,089,931 bytes prose out : 626,158 bytes words : 98,845 macros expanded: 3,174 citations : 1,109 (551 unique keys) PDF-only : 1 with no body to indexAll three runtimes produce the same 98,845 words.
The interesting decision is where to cut the job
Section titled “The interesting decision is where to cut the job”Indexing a paper is two steps that look very different. Unpacking the archive is I/O and decompression; parsing is string work. The obvious move is to offload the parsing, because that is the part that looks like computation.
In this benchmark, moving both steps to workers was faster. The following comparison measures all three options.
Expected output (4 cores / 8 threads, Bun 1.4, 16 papers per iteration):
word parity check: host=98,845 split=98,845 worker=98,845 OK match
LaTeX corpus benchmark (mitata)workload: gunzip + untar + inline inputs + expand macros + extract prosepapers per iteration: 16threads: 4
benchmark avg (min … max)host: unpack + parse 102.95 ms/iter (95.50 ms … 110.16 ms)worker: parse only, host unpacks 87.58 ms/iter (80.56 ms … 96.66 ms)worker: unpack + parse 61.66 ms/iter (51.08 ms … 81.98 ms)
summary worker: unpack + parse 1.42x faster than worker: parse only, host unpacks 1.67x faster than host: unpack + parseThe middle row is the one to look at. Offloading the parse — the part that looks like the
real work — bought 1.18x, because the decompression stayed on the thread you were trying
to free, and gunzipSync blocks it completely while it runs. Sending a path instead of
a payload moves the whole job across, and the gunzip parallelizes along with everything
else: 1.67x from the same pool, on the same corpus.
It is also the smaller piece of code. The host never opens an archive, so loadArchive is
called in exactly one place — inside the task.
// The shape worth copying: the job crosses the boundary as a path.export const indexArchive = task<string, PaperRecord>({ f: async (path: string) => parsePaperHost(await loadArchive(path)),});The boundary is cheaper than it looks
Section titled “The boundary is cheaper than it looks”Sending a path lets the worker handle decompression as well as parsing. Payload transfer was a relatively small cost in this workload, as two additional measurements show.
- Sending 64 papers’ worth of sources — 3.9 MiB of LaTeX — through a task that does nothing but count the files costs about 11 ms.
- Returning every paper’s extracted prose (626 KB) instead of returning counters only measured the same, within noise, in both directions across thread counts.
So the site’s usual advice — return compact summaries, not large payloads — is a rule about ratios, not about sizes. When each item costs milliseconds of CPU, a 40 KB result may add no measurable overhead. When each item costs microseconds, the same result dominates. Measure the ratio for your workload rather than assuming the payload is the problem.
The stage that refused to parallelize
Section titled “The stage that refused to parallelize”The first version of this example used DecompressionStream, which is the web-standard way
to gunzip and works on all three runtimes. In these measurements, adding workers did not make it faster.
Gunzipping the same sixteen archives inside the pool, changing only the thread count:
| threads | DecompressionStream | node:zlib gunzipSync |
|---|---|---|
| 1 | 43.9 ms | 60.8 ms |
| 2 | 39.3 ms | 58.2 ms |
| 4 | 37.1 ms | 32.2 ms |
| 8 | 39.2 ms | 27.0 ms |
DecompressionStream is flat. gunzipSync starts out slower and then scales, because it
runs on the worker that called it instead of handing the work to something shared.
Real submissions are messy
Section titled “Real submissions are messy”Most of the parser exists because of things real papers do:
- A submission can contain no LaTeX at all. One of these sixteen is a PDF with a stub
.texthat\includepdfs it. It parses to zero words, and the record sayspdfOnlyrather than quietly reporting a paper with no content. Adam is another, which is why it is not in the list. - A tarball can hold several
\begin{document}files. YOLOv3 ships the paper, a rebuttal, supplementary material, and a style demo. Picking the longest file gets you the supplementary material; the paper’s own entry point is 3 KB that\inputs everything else. The parser chooses the file that produces the largest document after resolving its inputs. - Every author invents macros. One paper in this corpus expands 1,844 of them. Skip expansion and your index fills up with tokens nobody wrote.
%starts a comment,\%does not, and papers are full of commented-out drafts.
The extracted text was checked against pdftotext output from arXiv’s own PDFs, minus the
bibliography: the parser lands between 0.84x and 1.17x of the reference word count across
spot-checked papers. It is an indexing-grade extraction, not a typesetter.
Do not expect this to scale past a few workers
Section titled “Do not expect this to scale past a few workers”Sixteen papers is a small corpus, and it shows. One worker performs about the same as the host. Adding workers helps clearly up to about four, and past that two sweeps on the same machine disagreed with each other — one kept improving to eight, the other got worse at six. That spread is larger than the effect, so there is no per-thread table here worth printing.
What is stable is the shape, and the reason for it. The sequential cost of the corpus is about 119 ms, and the most expensive single paper is 21 ms, so a perfect four-way split would finish in 30 ms. The pool lands closer to 55-65 ms. The corpus simply runs out: sixteen items across four workers is four items each, and the round ends when the last worker finishes.
Two practical lessons follow:
- Size the corpus to the pool, not the pool to the machine. If you have sixteen items, four workers were enough in these measurements. Try processing more items per round before adding workers.
- The largest item sets a floor you cannot cross. Here that floor is 21 ms and it is not yet binding. If your corpus has one item that is 10x the median, it will be, and no pool size fixes it — you would have to split that item.
import { mkdir, readFile, writeFile } from "node:fs/promises";import { gunzipSync } from "node:zlib";import { fileURLToPath } from "node:url";import { task } from "knitting";import { type PaperRecord, parsePaperHost, type PaperSource } from "./tex_parse.ts";
// Sixteen well-known papers, chosen for the spread: the smallest is 22 KB of// LaTeX and the largest is 226 KB. Two of them are PDF-only submissions, which// is a case any real corpus job has to survive.export const PAPER_IDS = [ "1706.03762", // Attention Is All You Need "1512.03385", // Deep Residual Learning "1409.1556", // VGG "1810.04805", // BERT "1406.2661", // Generative Adversarial Nets "1505.04597", // U-Net "1312.6114", // Auto-Encoding Variational Bayes "1502.03167", // Batch Normalization "1301.3781", // word2vec "1607.06450", // Layer Normalization "2010.11929", // An Image is Worth 16x16 Words "1503.02531", // Distilling the Knowledge in a Neural Network "2001.08361", // Scaling Laws for Neural Language Models "1910.10683", // T5 "1611.03530", // Rethinking Generalization "1804.02767", // YOLOv3];
const CACHE_DIR = fileURLToPath(new URL("./papers/", import.meta.url));
// arXiv asks that you not hammer the e-print endpoint. This downloads each// paper once, one at a time, and every later run reads the cache. If you want// thousands of papers, use arXiv's bulk access rather than a loop like this.const REQUEST_SPACING_MS = 3_000;
export function archivePath(id: string): string { return `${CACHE_DIR}${id}.tar.gz`;}
/** Download whatever is missing from the cache. Silent when there is nothing to do. */export async function ensureCorpus(ids: string[] = PAPER_IDS): Promise<string[]> { await mkdir(CACHE_DIR, { recursive: true }); let fetched = 0;
for (const id of ids) { const path = archivePath(id); try { await readFile(path); continue; } catch { // not cached yet } if (fetched === 0) console.log(`Fetching sources from arXiv into ${CACHE_DIR}`); else await new Promise((resolve) => setTimeout(resolve, REQUEST_SPACING_MS));
const response = await fetch(`https://arxiv.org/e-print/${id}`, { headers: { "user-agent": "knitting-docs-example/1.0" }, }); if (!response.ok) throw new Error(`arXiv ${id}: HTTP ${response.status}`); const bytes = new Uint8Array(await response.arrayBuffer()); await writeFile(path, bytes); fetched++; console.log(` ${id} ${bytes.byteLength.toLocaleString()} bytes`); }
if (fetched > 0) console.log(`Cached ${fetched} new archives.\n`); return ids.map(archivePath);}
/** * Read the .tex entries out of a tar archive. * * tar is 512-byte header blocks followed by padded contents, which is little * enough format to hand-roll rather than take a dependency on for one example. */function untarTexFiles(archive: Uint8Array): Record<string, string> { const decoder = new TextDecoder(); const files: Record<string, string> = {}; let offset = 0; let longName: string | null = null;
while (offset + 512 <= archive.length) { const header = archive.subarray(offset, offset + 512); if (header.every((byte) => byte === 0)) break;
const field = (start: number, length: number) => decoder.decode(header.subarray(start, start + length)) .replace(/\0.*$/, "").trim();
const prefix = field(345, 155); const rawName = field(0, 100); const name = longName ?? (prefix ? `${prefix}/${rawName}` : rawName); const size = Number.parseInt(field(124, 12), 8) || 0; const type = String.fromCharCode(header[156]!); const body = archive.subarray(offset + 512, offset + 512 + size); offset += 512 + Math.ceil(size / 512) * 512;
if (type === "L") { // GNU long name: this entry's content is the next header's real name. longName = decoder.decode(body).replace(/\0.*$/, ""); continue; } longName = null; if (type !== "0" && type !== "\0") continue; if (!/\.(tex|ltx)$/i.test(name)) continue; files[name] = decoder.decode(body); }
return files;}
/** Read one archive off disk and unpack the .tex files in it. */export async function loadArchive(path: string): Promise<PaperSource> { const id = path.split("/").pop()!.replace(/\.tar\.gz$/, ""); const archive = await readFile(path); // `node:zlib`, not `DecompressionStream`. Both work on all three runtimes and // the streaming one looks more idiomatic, but it does not get faster when you // add workers -- measured flat from 1 thread to 8. `gunzipSync` runs inside // the worker that called it, so it scales with the pool like everything else. return { id, files: untarTexFiles(gunzipSync(archive)) };}
/** * Read, unpack and parse one archive, start to finish, inside the worker. * * This is the shape worth copying. The job crosses the boundary as a path -- * a few dozen bytes -- rather than as the megabyte of LaTeX inside the archive, * and the gunzip goes parallel along with the parsing instead of staying on the * host as a serial prelude to it. */export const indexArchive = task<string, PaperRecord>({ f: async (path: string) => parsePaperHost(await loadArchive(path)),});/** One arXiv submission: every .tex file in the tarball, keyed by path. */export type PaperSource = { id: string; files: Record<string, string>;};
// Reading TeX far enough to find the prose in it. A real LaTeX engine this is// not -- it is the amount of the language you have to understand to turn a// submission into something you can index, which is a much smaller language// than the one TeX actually implements.
/** * Read a balanced `{...}` group starting at the opening brace. * * This is the workhorse of the whole parser, and the reason parsing a paper is * real CPU work rather than a few regexes: braces nest, and TeX lets you escape * them, so you cannot match them without walking the string. */export function readGroup( src: string, open: number,): { body: string; end: number } | null { if (src[open] !== "{") return null; let depth = 0; for (let i = open; i < src.length; i++) { const ch = src[i]; if (ch === "\\") { i++; // an escaped brace is a character, not a delimiter continue; } if (ch === "{") depth++; else if (ch === "}") { depth--; if (depth === 0) return { body: src.slice(open + 1, i), end: i + 1 }; } } return null;}
export function skipOptional(src: string, at: number): number { if (src[at] !== "[") return at; const close = src.indexOf("]", at); return close === -1 ? at : close + 1;}
export function readCommandName(src: string, backslash: number): string { let i = backslash + 1; while (i < src.length && /[A-Za-z]/.test(src[i]!)) i++; if (i === backslash + 1) return src[i] ?? ""; // \\, \%, \& and friends let name = src.slice(backslash + 1, i); if (src[i] === "*") name += "*"; return name;}
/** Drop `%` comments, but not `\%`, and keep the line structure. */export function stripComments(src: string): string { const lines = src.split("\n"); for (let i = 0; i < lines.length; i++) { const line = lines[i]!; let at = line.indexOf("%"); // A `%` is only a comment when the backslashes before it are an even run; // `\%` is a literal percent sign, and `\\%` is a line break then a comment. while (at > 0) { let backslashes = 0; for (let j = at - 1; j >= 0 && line.charCodeAt(j) === 92; j--) backslashes++; if (backslashes % 2 === 0) break; at = line.indexOf("%", at + 1); } if (at !== -1) lines[i] = line.slice(0, at); } return lines.join("\n");}
/** * Splice `\input`/`\include` files in, so the paper becomes one string. * * A paper is only meaningful as a whole -- a task that took one .tex file at a * time would be splitting a document mid-sentence. This is why the unit of work * here is a submission, not a file. */function inlineInputs( files: Record<string, string>, entry: string, seen: Set<string>,): string { if (seen.has(entry)) return ""; seen.add(entry); const src = stripComments(files[entry] ?? ""); let out = ""; let i = 0; while (i < src.length) { const backslash = src.indexOf("\\", i); if (backslash === -1) { out += src.slice(i); break; } out += src.slice(i, backslash); const name = readCommandName(src, backslash); if (name !== "input" && name !== "include") { out += src.slice(backslash, backslash + 1 + Math.max(name.length, 1)); i = backslash + 1 + Math.max(name.length, 1); continue; } const group = readGroup(src, backslash + 1 + name.length); if (!group) { out += src.slice(backslash, backslash + 1 + name.length); i = backslash + 1 + name.length; continue; } out += "\n" + inlineInputs(files, resolveName(files, group.body), seen) + "\n"; i = group.end; } return out;}
function resolveName(files: Record<string, string>, raw: string): string { const wanted = raw.trim().replace(/^\.\//, ""); for (const candidate of [wanted, `${wanted}.tex`]) { if (files[candidate] !== undefined) return candidate; } // Tarballs nest sources in directories; match on the basename as a fallback. const base = wanted.split("/").pop()!; for (const key of Object.keys(files)) { const keyBase = key.split("/").pop()!; if (keyBase === base || keyBase === `${base}.tex`) return key; } return wanted;}
type Macro = { arity: number; body: string };
/** * Collect `\newcommand`-style definitions. * * Every paper invents its own shorthand -- `\newcommand{\R}{\mathbb{R}}`, * `\newcommand{\model}{Transformer}` -- and leaving them unexpanded means your * index is full of tokens no reader ever typed. */export function collectMacros(src: string): Map<string, Macro> { const macros = new Map<string, Macro>(); const definers = /\\(?:re)?newcommand\*?|\\providecommand\*?|\\def/g; let match: RegExpExecArray | null; while ((match = definers.exec(src)) !== null) { let i = match.index + match[0].length; let name: string | null = null;
if (src[i] === "{") { const group = readGroup(src, i); if (!group) continue; const inner = group.body.trim(); if (!inner.startsWith("\\")) continue; name = inner.slice(1); i = group.end; } else if (src[i] === "\\") { name = readCommandName(src, i); i += 1 + name.length; } if (!name || !/^[A-Za-z]+\*?$/.test(name)) continue;
let arity = 0; if (src[i] === "[") { const close = src.indexOf("]", i); if (close !== -1) { arity = Number.parseInt(src.slice(i + 1, close), 10) || 0; i = close + 1; } } i = skipOptional(src, i); // an optional-argument default value const body = readGroup(src, i); if (!body) continue; macros.set(name, { arity, body: body.body }); } return macros;}
// Macro bodies call other macros, so one pass is not enough. Papers do not nest// them deeply, though, and a fixed ceiling is what keeps a recursive definition// from turning one bad submission into a hung worker.const MAX_EXPANSION_PASSES = 4;
export function expandMacros( src: string, macros: Map<string, Macro>,): { text: string; expansions: number } { let text = src; let expansions = 0;
for (let pass = 0; pass < MAX_EXPANSION_PASSES; pass++) { let out = ""; let i = 0; let hits = 0;
while (i < text.length) { const backslash = text.indexOf("\\", i); if (backslash === -1) { out += text.slice(i); break; } out += text.slice(i, backslash); const name = readCommandName(text, backslash); const macro = name ? macros.get(name) : undefined; if (!macro) { const width = 1 + Math.max(name.length, 1); out += text.slice(backslash, backslash + width); i = backslash + width; continue; }
let cursor = backslash + 1 + name.length; const args: string[] = []; for (let a = 0; a < macro.arity; a++) { while (text[cursor] === " ") cursor++; const group = readGroup(text, cursor); if (!group) break; args.push(group.body); cursor = group.end; } if (args.length < macro.arity) { const width = 1 + name.length; out += text.slice(backslash, backslash + width); i = backslash + width; continue; }
out += macro.body.replace(/#(\d)/g, (_, digit: string) => args[Number(digit) - 1] ?? ""); i = cursor; hits++; }
text = out; expansions += hits; if (hits === 0) break; }
return { text, expansions };}
/** * Pick the file that starts the document, and return it already resolved. * * A tarball has no manifest, so the entry point is whatever holds * `\begin{document}` -- but submissions routinely ship several: a paper, a * rebuttal, a supplementary, a leftover style demo, each with its own * `\documentclass`. The one you want is not the longest file, it is the one * that pulls in the most document once its `\input`s are resolved. YOLOv3 is * the case that forces this: its entry point is 3 KB that inlines the paper, * while the supplementary next to it is 11 KB on its own. * * Resolving is the expensive part, so the winner's text comes back with it * rather than being rebuilt by the caller. */export function resolveDocument( files: Record<string, string>,): { main: string; merged: string } | null { let best: { main: string; merged: string } | null = null; for (const [name, text] of Object.entries(files)) { if (!text.includes("\\begin{document}")) continue; const merged = inlineInputs(files, name, new Set()); if (best === null || merged.length > best.merged.length) { best = { main: name, merged }; } } return best;}
/** Everything between `\begin{document}` and `\end{document}`. */export function documentBody(src: string): string { const start = src.indexOf("\\begin{document}"); if (start === -1) return src; const from = start + "\\begin{document}".length; const end = src.lastIndexOf("\\end{document}"); return end > from ? src.slice(from, end) : src.slice(from);}import { task } from "knitting";import { documentBody, type PaperSource, readGroup, resolveDocument, skipOptional, readCommandName, collectMacros, expandMacros,} from "./tex_scan.ts";
export type { PaperSource };
export type Section = { level: number; title: string;};
export type PaperRecord = { id: string; title: string; sections: Section[]; citationKeys: string[]; /** Comments stripped, macros expanded, math replaced with placeholders. */ text: string; texBytes: number; textBytes: number; words: number; equations: number; figures: number; tables: number; macrosDefined: number; macrosExpanded: number; /** True when the tarball is a PDF wrapped in a stub .tex, with no real body. */ pdfOnly: boolean;};
// A real LaTeX engine this is not. It is the amount of LaTeX you have to// understand to turn a paper into something you can index or embed, which is a// much smaller language than the one TeX actually implements.
const SECTION_COMMANDS: Record<string, number> = { chapter: 0, section: 1, subsection: 2, subsubsection: 3,};
// Macros whose braces hold prose: drop the macro, keep the argument.const UNWRAP = new Set([ "emph", "textit", "textbf", "texttt", "textsc", "textrm", "textsf", "underline", "mbox", "text", "mathrm", "footnote", "caption", "title", "author", "abstract", "paragraph", "subparagraph", "textsuperscript",]);
// Macros whose braces hold machinery: drop the macro and the argument with it.const DROP_WITH_ARG = new Set([ "documentclass", "usepackage", "bibliography", "bibliographystyle", "includegraphics", "label", "ref", "eqref", "pageref", "cite", "citep", "citet", "citealp", "citeauthor", "citeyear", "input", "include", "hypersetup", "setlength", "vspace", "hspace", "includepdf", "pdfoutput", "newtheorem", "geometry", "definecolor", "url", "thanks", "nocite", "color", "pagestyle", "bibliographyfont", "addtolength",]);
// `\href{url}{text}`, `\textcolor{red}{text}`: two groups, and only the last// one is prose.const KEEP_LAST_GROUP = new Set(["href", "textcolor", "hyperref", "colorbox"]);
// `\newcommand{\dmodel}{d_{\text{model}}}` has two groups and neither is prose.// Treated like the others it would paste every macro body into the output, which// is how you end up indexing `d_\textmodel` as if an author had written it.const DEFINERS = new Set([ "newcommand", "renewcommand", "providecommand", "def", "DeclareMathOperator",]);
// Environments whose content is not prose at all.const SKIP_ENVIRONMENTS = new Set([ "equation", "equation*", "align", "align*", "eqnarray", "eqnarray*", "gather", "gather*", "multline", "multline*", "displaymath", "math", "array", "tabular", "tabular*", "tabularx", "matrix", "pmatrix", "bmatrix", "verbatim", "lstlisting", "tikzpicture", "thebibliography", "algorithmic", "algorithm", "picture", "filecontents", "filecontents*",]);
const MATH_ENVIRONMENTS = new Set([ "equation", "equation*", "align", "align*", "eqnarray", "eqnarray*", "gather", "gather*", "multline", "multline*", "displaymath",]);
type EnvCounts = { equations: number; figures: number; tables: number };
/** * Remove environments whose content is not prose, counting them on the way out. * * Nesting is the whole difficulty: a `figure` holding a `tabular` holding an * `align` has to come out as one unit, so this tracks depth per environment * name rather than searching for the next `\end`. */function stripEnvironments(src: string): { text: string; counts: EnvCounts } { const counts: EnvCounts = { equations: 0, figures: 0, tables: 0 }; const marker = /\\(begin|end)\s*\{([^}]*)\}/g; let out = ""; let cursor = 0; let skipping: string | null = null; let depth = 0; let match: RegExpExecArray | null;
while ((match = marker.exec(src)) !== null) { const [, kind, rawName] = match; const name = rawName!.trim();
if (skipping === null) { if (kind === "begin") { const base = name.replace(/\*$/, ""); if (MATH_ENVIRONMENTS.has(name)) counts.equations++; else if (base === "figure" || base === "subfigure") counts.figures++; else if (base === "table") counts.tables++;
if (SKIP_ENVIRONMENTS.has(name)) { out += src.slice(cursor, match.index) + "\n"; skipping = name; depth = 1; continue; } } out += src.slice(cursor, match.index); cursor = marker.lastIndex; continue; }
if (name !== skipping) continue; if (kind === "begin") depth++; else if (--depth === 0) { skipping = null; cursor = marker.lastIndex; } }
if (skipping === null) out += src.slice(cursor); return { text: out, counts };}
function stripMath(src: string): { text: string; equations: number } { let equations = 0; const bump = () => { equations++; return " "; }; const text = src .replace(/\$\$[\s\S]*?\$\$/g, bump) .replace(/\\\[[\s\S]*?\\\]/g, bump) .replace(/(?<!\\)\$(?:\\.|[^$\\])*\$/g, () => " ") .replace(/\\\((?:[\s\S]*?)\\\)/g, () => " "); return { text, equations };}
type Walked = { text: string; sections: Section[]; citationKeys: string[];};
/** * The last pass: keep prose, drop markup, and pick up structure on the way. * * Three kinds of macro get three different treatments -- some hold prose worth * keeping, some hold machinery worth dropping whole, and the long tail is * neither, so the macro goes and its braces stay. */function walkMacros( src: string, sections: Section[] = [], citationKeys: string[] = [], depth = 0,): Walked { let out = ""; let i = 0;
while (i < src.length) { const backslash = src.indexOf("\\", i); if (backslash === -1) { out += src.slice(i); break; } out += src.slice(i, backslash);
const name = readCommandName(src, backslash); if (!name) { i = backslash + 1; continue; } let cursor = backslash + 1 + name.length; const base = name.replace(/\*$/, "");
if (DEFINERS.has(base)) { // Drop the name group and the body group both. if (src[cursor] === "\\") cursor += 1 + readCommandName(src, cursor).length; let group = readGroup(src, cursor); if (group) cursor = group.end; cursor = skipOptional(src, skipOptional(src, cursor)); group = readGroup(src, cursor); if (group) cursor = group.end; out += " "; i = cursor; continue; }
if (base in SECTION_COMMANDS) { cursor = skipOptional(src, cursor); const group = readGroup(src, cursor); if (group) { const title = cleanInline(group.body); sections.push({ level: SECTION_COMMANDS[base]!, title }); out += `\n\n${title}\n\n`; i = group.end; continue; } }
if (base.startsWith("cite") || base === "nocite") { cursor = skipOptional(src, skipOptional(src, cursor)); const group = readGroup(src, cursor); if (group) { for (const key of group.body.split(",")) { const trimmed = key.trim(); if (trimmed) citationKeys.push(trimmed); } out += " "; i = group.end; continue; } }
if (KEEP_LAST_GROUP.has(base)) { const url = readGroup(src, skipOptional(src, cursor)); const label = url ? readGroup(src, url.end) : null; if (label) { out += walkMacros(label.body, sections, citationKeys, depth + 1).text; i = label.end; continue; } }
if (UNWRAP.has(base)) { cursor = skipOptional(src, cursor); const group = readGroup(src, cursor); // The braces held prose, and prose holds more macros -- `\emph{\texttt{x}}` // only comes out clean if the body is walked too, not pasted through. if (group && depth < MAX_UNWRAP_DEPTH) { out += walkMacros(group.body, sections, citationKeys, depth + 1).text; i = group.end; continue; } }
if (DROP_WITH_ARG.has(base)) { cursor = skipOptional(src, cursor); const group = readGroup(src, cursor); if (group) { out += " "; i = group.end; continue; } }
// Everything else: drop the command, keep whatever it wrapped. out += " "; i = cursor; }
return { text: out, sections, citationKeys };}
// Prose nests, but not forever. The ceiling keeps a pathological document from// turning into a stack overflow inside a worker.const MAX_UNWRAP_DEPTH = 12;
function cleanInline(raw: string): string { return raw .replace(/\\[A-Za-z]+\*?/g, " ") .replace(/\\./g, " ") // \\, \&, \% and the rest of the escapes .replace(/[{}$]/g, "") .replace(/\s+/g, " ") .trim();}
function tidy(raw: string): string { return raw .replace(/[{}]/g, "") .replace(/~/g, " ") .replace(/[ \t]+/g, " ") .replace(/ ?\n ?/g, "\n") .replace(/\n{3,}/g, "\n\n") .trim();}
export function parsePaperHost(source: PaperSource): PaperRecord { const texBytes = Object.values(source.files) .reduce((total, text) => total + text.length, 0); const resolved = resolveDocument(source.files);
if (resolved === null) { return emptyRecord(source.id, texBytes); }
const merged = resolved.merged; // Definitions live in the preamble; prose lives in the body. Collect from the // whole file, then index only what is between \begin{document} and its \end, // because a preamble full of \usepackage lines is not text anybody wrote. const macros = collectMacros(merged); const { text: expanded, expansions } = expandMacros(documentBody(merged), macros);
const titleMatch = /\\title\s*(?:\[[^\]]*\])?\s*\{/.exec(merged); const titleGroup = titleMatch ? readGroup(merged, titleMatch.index + titleMatch[0].length - 1) : null;
const { text: withoutEnvs, counts } = stripEnvironments(expanded); const { text: withoutMath, equations: inlineEquations } = stripMath( withoutEnvs, ); const walked = walkMacros(withoutMath); const text = tidy(walked.text); const words = text.length === 0 ? 0 : text.split(/\s+/).length;
return { id: source.id, title: titleGroup ? cleanInline(titleGroup.body) : "", sections: walked.sections, citationKeys: walked.citationKeys, text, texBytes, textBytes: text.length, words, equations: counts.equations + inlineEquations, figures: counts.figures, tables: counts.tables, macrosDefined: macros.size, macrosExpanded: expansions, // Some submissions are a PDF with a two-line .tex wrapper around it. There // is nothing to index, and a corpus job should say so rather than record a // paper with no words in it. pdfOnly: merged.includes("\\includepdf") && words < 200, };}
function emptyRecord(id: string, texBytes: number): PaperRecord { return { id, title: "", sections: [], citationKeys: [], text: "", texBytes, textBytes: 0, words: 0, equations: 0, figures: 0, tables: 0, macrosDefined: 0, macrosExpanded: 0, pdfOnly: false, };}
export const parsePaper = task<PaperSource, PaperRecord>({ f: parsePaperHost });import { createPool, isMain } from "knitting";import { ensureCorpus, indexArchive } from "./arxiv_corpus.ts";import type { PaperRecord } from "./tex_parse.ts";
const THREADS = 4;
async function main() { // First run downloads the sources from arXiv; later runs read the cache. const paths = await ensureCorpus(); using pool = createPool({ threads: THREADS })({ indexArchive });
const started = performance.now(); const records = await Promise.all(paths.map(pool.call.indexArchive)); const elapsedMs = performance.now() - started;
const total = (pick: (record: PaperRecord) => number) => records.reduce((sum, record) => sum + pick(record), 0); const citationKeys = new Set(records.flatMap((record) => record.citationKeys));
console.log( `\nIndexed ${records.length} arXiv submissions on ${THREADS} workers ` + `in ${elapsedMs.toFixed(0)} ms\n`, ); console.log(" id words sec eq fig tab cites title"); for (const record of [...records].sort((a, b) => b.words - a.words)) { console.log( ` ${record.id.padEnd(11)}` + `${record.words.toLocaleString().padStart(6)}` + `${String(record.sections.length).padStart(6)}` + `${String(record.equations).padStart(5)}` + `${String(record.figures).padStart(5)}` + `${String(record.tables).padStart(5)}` + `${String(record.citationKeys.length).padStart(7)} ` + (record.pdfOnly ? "(PDF-only submission)" : record.title.slice(0, 44)), ); }
console.log("\n LaTeX in :", total((r) => r.texBytes).toLocaleString(), "bytes"); console.log(" prose out :", total((r) => r.textBytes).toLocaleString(), "bytes"); console.log(" words :", total((r) => r.words).toLocaleString()); console.log(" macros expanded:", total((r) => r.macrosExpanded).toLocaleString()); console.log( " citations :", `${total((r) => r.citationKeys.length).toLocaleString()} ` + `(${citationKeys.size.toLocaleString()} unique keys)`, ); console.log(" PDF-only :", records.filter((r) => r.pdfOnly).length, "with no body to index");}
if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; });}import { createPool, isMain } from "knitting";import { bench, boxplot, run, summary } from "mitata";import { ensureCorpus, indexArchive, loadArchive } from "./arxiv_corpus.ts";import { parsePaper, parsePaperHost } from "./tex_parse.ts";
const THREADS = 4;
async function main() { const paths = await ensureCorpus(); using pool = createPool({ threads: THREADS })({ indexArchive, parsePaper });
const onHost = () => hostJob(paths); const splitJob = () => hostUnpacksWorkerParses(pool.call.parsePaper, paths); const workerJob = () => workerDoesBoth(pool.call.indexArchive, paths);
const hostWords = await onHost(); const splitWords = await splitJob(); const workerWords = await workerJob(); const parity = hostWords === splitWords && hostWords === workerWords; console.log( `word parity check: host=${hostWords.toLocaleString()} ` + `split=${splitWords.toLocaleString()} ` + `worker=${workerWords.toLocaleString()} ` + (parity ? "OK match" : "MISMATCH"), ); if (!parity) throw new Error("Word counts differ between placements.");
console.log("\nLaTeX corpus benchmark (mitata)"); console.log("workload: gunzip + untar + inline inputs + expand macros + extract prose"); console.log("papers per iteration:", paths.length); console.log("threads:", THREADS, "\n");
let sink = 0; boxplot(() => { summary(() => { bench("host: unpack + parse", async () => { sink = await onHost(); });
bench("worker: parse only, host unpacks", async () => { sink = await splitJob(); });
bench("worker: unpack + parse", async () => { sink = await workerJob(); }); }); });
await run(); console.log("last word count:", sink.toLocaleString());}
/** Everything on the main thread. */async function hostJob(paths: string[]): Promise<number> { let words = 0; for (const path of paths) { words += parsePaperHost(await loadArchive(path)).words; } return words;}
/** * Split down the middle: the host gunzips and untars, the worker parses. * * This is the tempting shape -- offload the part that looks expensive -- and it * leaves the decompression on the thread you were trying to free. */async function hostUnpacksWorkerParses( call: (source: Awaited<ReturnType<typeof loadArchive>>) => Promise<{ words: number }>, paths: string[],): Promise<number> { const sources = await Promise.all(paths.map(loadArchive)); const records = await Promise.all(sources.map(call)); return records.reduce((total, record) => total + record.words, 0);}
/** The whole job on the worker: only a path crosses the boundary. */async function workerDoesBoth( call: (path: string) => Promise<{ words: number }>, paths: string[],): Promise<number> { const records = await Promise.all(paths.map(call)); return records.reduce((total, record) => total + record.words, 0);}
if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; });}When this matters
Section titled “When this matters”This pattern applies to document ingestion: files on disk or in a bucket that need decompression, parsing, and text extraction before further processing. Search indexes, RAG pipelines and dataset builds all look like this, and the per-item cost is high enough that doing it on the request thread is not an option.
The pattern to take away is the task signature. indexArchive takes a path and returns a
record, which means the pool owns the whole pipeline for one item — reading included. That
is both faster than splitting the pipeline across the boundary and simpler to write.