LLM input shaping
This example trims LLM prompts to fit a token budget before an API call. It
counts tokens with tiktoken, drops the oldest conversation turns first, and
shortens the query only if removing every turn is still not enough. The budgeting logic is model-agnostic — the example uses the tokenizer for gpt-4o-mini, but the pattern is the same for any model with an input limit.
How it works
Section titled “How it works”- The host builds synthetic chat traffic: same system prefix, 3 to 14 turns of history, and a query.
- Each task builds the full prompt and counts tokens with
tiktoken. - If the prompt is over budget, it drops the oldest turns one at a time, re-counting after each.
- If no turns are left and it is still over, it clips the query to whatever budget remains.
- The host aggregates token savings and trim counts.
Four files:
token_budget.ts— the budgeting logic and the two tasksprompt_fixtures.ts— the synthetic conversations, so both scripts share one generatorrun_prompt_token_budget.ts— what budgeting actually does to a workloadbench_prompt_token_budget.ts— host vs workers, withmitata
There are two tasks because the right return shape depends on what you need. preparePrompt
returns the whole PromptPlan, prompt string included, which is what an application wants.
summarizeBatch takes a batch and returns counters only — no prompt strings — which is what
you want when you are measuring, or when you only need the accounting.
Example budget decision
Section titled “Example budget decision”Minimal usage:
using pool = createPool({ threads: 3 })({ preparePrompt });const plan = await pool.call.preparePrompt({ model: "gpt-4o-mini", systemPrefix: "You are a docs assistant.", history: ["Need guidance on schema validation.", "Keep the answer short."], query: "Give a migration plan and one code example.", maxInputTokens: 400,});Output shape:
{ prompt: "...final prompt string...", rawInputTokens: 1061, inputTokens: 399, staticTokens: 47, dynamicTokens: 352, trimmedTurns: 3, queryWasTrimmed: true,}The useful part is not just the final prompt. You also get the bookkeeping needed to explain
why a prompt was trimmed and by how much — which turns were dropped, whether the query itself
had to be clipped, and how much of the budget the fixed prefix uses.
bun src/run_prompt_token_budget.tsdeno run -A src/run_prompt_token_budget.tsnpx tsx src/run_prompt_token_budget.tsExpected output:
Prompt token budgeting requests : 2,000 budget : 400 tokens per prompt raw tokens : 904,702 budgeted tokens : 684,035 saved : 220,667 (24.4%) trimmed : 1,057 / 2,000 requests turns dropped : 4,286 query clipped : 118 elapsed : 1640 ms
One request, in detail: 1061 -> 399 tokens 3 turns dropped query clipped: true 47 of those tokens are the fixed prefixChoose a budget that some sample prompts exceed. Otherwise, the example only measures token counting and never exercises the trimming logic.
Optional benchmark
Section titled “Optional benchmark”The benchmark compares budgeting throughput on the host with the same work on workers. Batches are sent
through summarizeBatch, so only counters come back, never prompt strings.
token parity check: host=341,835 worker=341,835 OK match
Prompt token budgeting benchmark (mitata)workload: build prompt + tokenize + trim to budgetrequests per iteration: 1,000budget: 400 tokenstrimmed: 528 requests, 24.4% of tokens savedthreads: 3 | batch: 32
benchmark avg (min … max)host (1,000 req) 1.45 s/iter (1.39 s … 1.46 s)knitting (3 threads, 1,000 req) 471.93 ms/iter (468.14 ms … 497.00 ms)
summary knitting (3 threads, 1,000 req) 3.08x faster than host (1,000 req)Why more workers make this slower
Section titled “Why more workers make this slower”This workload stops getting faster after a few workers. The measurements below show why choosing a pool size requires more than counting cores.
On a 4-core laptop the best setting measured here was three workers plus the
inliner — four lanes for four cores — at roughly 3x. Beyond that, performance drops: eight workers measured consistently slower than three or four. One
worker performed about the same as the host, within measurement noise.
With plain arithmetic in place of tokenization, the same pool and batching
reached 4.8x at eight threads on the same laptop. This suggests that the
tokenizer is limiting scaling in this workload.
The reason is that tiktoken is a WASM module with a large BPE table, and every
worker instantiates its own copy. Resident memory here went from 67 MB at
baseline to 630 MB with eight workers — roughly 55-60 MB per worker. Those
tables are not shared and exceed the available CPU cache. Adding workers can
therefore increase cache pressure without improving throughput.
Two practical consequences:
- Size the pool for the tokenizer, not for
os.cpus().length. Start at 2-4 lanes and measure. The inliner helped in this benchmark — it adds a lane without adding another encoder-sized worker. - Keep many more batches than lanes. Batching amortizes dispatch, but one
giant batch per lane cannot rebalance. At 1,000 requests in batches of 250 the
pool measured slower than the host (
0.81x); the same work in batches of 32 was the fastest setting tried.
import { task } from "knitting";import { encoding_for_model } from "tiktoken";
export type PromptInput = { model: string; systemPrefix: string; history: string[]; query: string; maxInputTokens: number;};
export type PromptPlan = { prompt: string; rawInputTokens: number; inputTokens: number; staticTokens: number; dynamicTokens: number; trimmedTurns: number; queryWasTrimmed: boolean;};
/** What a batch of plans adds up to. No prompt strings, so it is cheap to return. */export type PromptBudgetSummary = { rawTokens: number; budgetedTokens: number; staticTokens: number; dynamicTokens: number; trimmedRuns: number; queryTrimmedRuns: number; turnsDropped: number;};
type Encoder = ReturnType<typeof encoding_for_model>;
const decoder = new TextDecoder();const encoderCache = new Map<string, Encoder>();const staticTokenCache = new Map<string, number>();
// A tiktoken encoder is a WASM instance holding its own BPE table, and it costs// tens of megabytes resident. This cache lives per worker, so the real ceiling// is MAX_ENCODERS * threads. Keep it small, and free whatever you evict.const MAX_ENCODERS = 2;
function getEncoder(model: string): Encoder { const cached = encoderCache.get(model); if (cached) return cached;
const enc = encoding_for_model(model as never); encoderCache.set(model, enc);
if (encoderCache.size > MAX_ENCODERS) { const oldest = encoderCache.keys().next().value!; encoderCache.get(oldest)!.free(); encoderCache.delete(oldest); }
return enc;}
/** The system prefix is identical on every request, so tokenize it once. */function getStaticTokens(model: string, prefix: string, enc: Encoder): number { const key = `${model}\x1f${prefix}`; const cached = staticTokenCache.get(key); if (cached !== undefined) return cached;
const value = enc.encode(prefix).length; staticTokenCache.set(key, value); return value;}
export function clearPromptBudgetCaches(): void { for (const enc of encoderCache.values()) enc.free(); encoderCache.clear(); staticTokenCache.clear();}
function normalizeText(value: string): string { return value.replace(/\s+/g, " ").trim();}
function buildPrompt(prefix: string, history: string[], query: string): string { const rows = [prefix.trim(), "", "Conversation context:"]; for (let i = 0; i < history.length; i++) { rows.push(`- Turn ${i + 1}: ${history[i]}`); } rows.push("", `User request: ${query}`); return rows.join("\n");}
function truncateToTokenBudget( enc: Encoder, text: string, maxTokens: number,): string { if (maxTokens <= 0) return "";
const tokens = enc.encode(text); if (tokens.length <= maxTokens) return text; return decoder.decode(enc.decode(tokens.slice(0, maxTokens)));}
/** * Fit a prompt inside `maxInputTokens`: drop the oldest turns first, and only * clip the query if dropping every turn still is not enough. */export function preparePromptHost(input: PromptInput): PromptPlan { const maxInputTokens = Math.max(64, input.maxInputTokens); const history = input.history.map(normalizeText).filter(Boolean); const enc = getEncoder(input.model); const staticTokens = getStaticTokens(input.model, input.systemPrefix, enc);
let query = normalizeText(input.query); let prompt = buildPrompt(input.systemPrefix, history, query); const rawInputTokens = enc.encode(prompt).length; let inputTokens = rawInputTokens; let trimmedTurns = 0; let queryWasTrimmed = false;
while (inputTokens > maxInputTokens && history.length > 0) { history.shift(); trimmedTurns++; prompt = buildPrompt(input.systemPrefix, history, query); inputTokens = enc.encode(prompt).length; }
// No turns left to drop and still over: the query itself is the problem. if (inputTokens > maxInputTokens) { const scaffolding = buildPrompt(input.systemPrefix, history, ""); const remaining = Math.max( 16, maxInputTokens - enc.encode(scaffolding).length, ); const clipped = truncateToTokenBudget(enc, query, remaining); queryWasTrimmed = clipped.length < query.length; query = clipped; prompt = buildPrompt(input.systemPrefix, history, query); inputTokens = enc.encode(prompt).length; }
return { prompt, rawInputTokens, inputTokens, staticTokens, dynamicTokens: Math.max(0, inputTokens - staticTokens), trimmedTurns, queryWasTrimmed, };}
/** * Budget a whole batch and return only the counters. Batching amortizes the * per-call dispatch, and dropping the prompt strings keeps the return small. */export function summarizeBatchHost( inputs: PromptInput[],): PromptBudgetSummary { const totals = emptySummary();
for (let i = 0; i < inputs.length; i++) { const plan = preparePromptHost(inputs[i]!); totals.rawTokens += plan.rawInputTokens; totals.budgetedTokens += plan.inputTokens; totals.staticTokens += plan.staticTokens; totals.dynamicTokens += plan.dynamicTokens; totals.turnsDropped += plan.trimmedTurns; if (plan.trimmedTurns > 0) totals.trimmedRuns++; if (plan.queryWasTrimmed) totals.queryTrimmedRuns++; }
return totals;}
export function emptySummary(): PromptBudgetSummary { return { rawTokens: 0, budgetedTokens: 0, staticTokens: 0, dynamicTokens: 0, trimmedRuns: 0, queryTrimmedRuns: 0, turnsDropped: 0, };}
export function mergeSummaries( parts: PromptBudgetSummary[],): PromptBudgetSummary { return parts.reduce((a, b) => ({ rawTokens: a.rawTokens + b.rawTokens, budgetedTokens: a.budgetedTokens + b.budgetedTokens, staticTokens: a.staticTokens + b.staticTokens, dynamicTokens: a.dynamicTokens + b.dynamicTokens, trimmedRuns: a.trimmedRuns + b.trimmedRuns, queryTrimmedRuns: a.queryTrimmedRuns + b.queryTrimmedRuns, turnsDropped: a.turnsDropped + b.turnsDropped, }), emptySummary());}
/** Returns the full plan, prompt string included. */export const preparePrompt = task<PromptInput, PromptPlan>({ f: preparePromptHost,});
/** Returns counters only. This is the one worth benchmarking. */export const summarizeBatch = task<PromptInput[], PromptBudgetSummary>({ f: summarizeBatchHost,});import type { PromptInput } from "./token_budget.ts";
/** * Synthetic chat traffic for the two scripts. Nothing here is part of the * budgeting logic -- it just produces conversations long enough that a real * budget has to trim them. */
export const SYSTEM_PREFIX = [ "You are a documentation assistant for a multi-threading library.", "Prefer concrete, short answers grounded in the provided context.", "If the data needed to answer is missing, say so directly.", "Never invent API surface that was not shown to you.",].join("\n");
const TOPICS = [ "token budgeting", "prompt caching", "parallel workers", "schema validation", "rendering pipelines", "markdown output", "compression tradeoffs", "latency under load", "shared memory buffers", "work stealing",];
const DETAIL = [ "Walk me through the tradeoffs before recommending anything.", "Assume a Node 22 service handling a few thousand requests a minute.", "We already tried the naive version and it pinned one core.", "Include the failure mode we should watch for in production.", "Keep the code sample under twenty lines if you can.",];
function pick<T>(values: T[], i: number): T { return values[i % values.length]!;}
export function buildPromptInputs( count: number, maxInputTokens: number, model = "gpt-4o-mini",): PromptInput[] { const inputs = new Array<PromptInput>(count);
for (let i = 0; i < count; i++) { const turns = 3 + (i % 12); const history = new Array<string>(turns); for (let t = 0; t < turns; t++) { history[t] = `Turn about ${pick(TOPICS, i + t)}: ${ pick(DETAIL, i + t * 3) } Earlier we settled on ${ pick(TOPICS, i + t + 5) }, so keep that decision in mind.`; }
const parts = [ `Compare ${pick(TOPICS, i)} with ${pick(TOPICS, i + 4)} for our workload.`, pick(DETAIL, i), "Give a short recommendation, a migration path, and the one metric that tells us it worked.", ];
// Every 17th user pastes a wall of log output. One of those can blow the // budget on its own, which is the only case where trimming reaches the query. if (i % 17 === 0) { for (let k = 0; k < 40; k++) { parts.push( `[worker ${k % 8}] task=${pick(TOPICS, i + k)} queued=${ k * 13 } claimed=${k * 7} elapsed_ms=${(k * 31) % 97}.${k % 10}`, ); } }
inputs[i] = { model, systemPrefix: SYSTEM_PREFIX, history, query: parts.join(" "), maxInputTokens, }; }
return inputs;}
export function batched<T>(values: T[], size: number): T[][] { const batches: T[][] = []; for (let i = 0; i < values.length; i += size) { batches.push(values.slice(i, i + size)); } return batches;}import { createPool, isMain } from "knitting";import { buildPromptInputs } from "./prompt_fixtures.ts";import { preparePrompt } from "./token_budget.ts";
// A budget only teaches you anything when it actually bites. These// conversations run roughly 200-1,450 tokens, so 400 trims about half of them.const REQUESTS = 2_000;const MAX_INPUT_TOKENS = 400;const THREADS = 3;
async function main() { const inputs = buildPromptInputs(REQUESTS, MAX_INPUT_TOKENS); using pool = createPool({ threads: THREADS })({ preparePrompt });
const started = performance.now(); const plans = await Promise.all(inputs.map(pool.call.preparePrompt)); const elapsedMs = performance.now() - started;
let rawTokens = 0; let budgetedTokens = 0; let trimmedRuns = 0; let queryTrimmedRuns = 0; let turnsDropped = 0;
for (const plan of plans) { rawTokens += plan.rawInputTokens; budgetedTokens += plan.inputTokens; turnsDropped += plan.trimmedTurns; if (plan.trimmedTurns > 0) trimmedRuns++; if (plan.queryWasTrimmed) queryTrimmedRuns++; }
const saved = rawTokens - budgetedTokens; const pct = (part: number) => `${((part / rawTokens) * 100).toFixed(1)}%`;
console.log("Prompt token budgeting"); console.log(" requests :", REQUESTS.toLocaleString()); console.log(" budget :", MAX_INPUT_TOKENS, "tokens per prompt"); console.log(" raw tokens :", rawTokens.toLocaleString()); console.log(" budgeted tokens :", budgetedTokens.toLocaleString()); console.log( " saved :", `${saved.toLocaleString()} (${pct(saved)})`, ); console.log( " trimmed :", `${trimmedRuns.toLocaleString()} / ${REQUESTS.toLocaleString()} requests`, ); console.log(" turns dropped :", turnsDropped.toLocaleString()); console.log(" query clipped :", queryTrimmedRuns.toLocaleString()); console.log(" elapsed :", `${elapsedMs.toFixed(0)} ms`);
// Every plan carries the bookkeeping needed to explain the decision. const example = plans.find((plan) => plan.queryWasTrimmed) ?? plans[0]!; console.log("\nOne request, in detail:"); console.log(` ${example.rawInputTokens} -> ${example.inputTokens} tokens`); console.log(` ${example.trimmedTurns} turns dropped`); console.log(` query clipped: ${example.queryWasTrimmed}`); console.log(` ${example.staticTokens} of those tokens are the fixed prefix`);}
if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; });}import { createPool, isMain } from "knitting";import { bench, boxplot, run, summary } from "mitata";import { batched, buildPromptInputs } from "./prompt_fixtures.ts";import { mergeSummaries, type PromptBudgetSummary, type PromptInput, summarizeBatch, summarizeBatchHost,} from "./token_budget.ts";
const REQUESTS = 1_000;const MAX_INPUT_TOKENS = 400;
// Three, not eight. Tokenizing does not scale the way plain compute does --// read "Why more workers make this slower" on the docs page before raising it. The inliner// earns its place here: it measured faster than a fourth worker.const THREADS = 3;
// Batch so there are many more batches than lanes. One batch per lane cannot// rebalance, and the slowest batch then holds up the whole round.const BATCH = 32;
async function main() { const inputs = buildPromptInputs(REQUESTS, MAX_INPUT_TOKENS); const batches = batched(inputs, BATCH); using pool = createPool({ threads: THREADS, inliner: { batchSize: 8 }, })({ summarizeBatch }); const callBatch = pool.call.summarizeBatch;
const hostTotals = runHost(batches); const workerTotals = await runWorkers(callBatch, batches); const parity = hostTotals.budgetedTokens === workerTotals.budgetedTokens && hostTotals.trimmedRuns === workerTotals.trimmedRuns; console.log( `token parity check: host=${hostTotals.budgetedTokens.toLocaleString()} ` + `worker=${workerTotals.budgetedTokens.toLocaleString()} ` + (parity ? "OK match" : "MISMATCH"), ); if (!parity) throw new Error("Host and worker budget totals differ.");
const saved = hostTotals.rawTokens - hostTotals.budgetedTokens; console.log("\nPrompt token budgeting benchmark (mitata)"); console.log("workload: build prompt + tokenize + trim to budget"); console.log("requests per iteration:", REQUESTS.toLocaleString()); console.log("budget:", MAX_INPUT_TOKENS, "tokens"); console.log( "trimmed:", `${hostTotals.trimmedRuns.toLocaleString()} requests, ` + `${((saved / hostTotals.rawTokens) * 100).toFixed(1)}% of tokens saved`, ); console.log("threads:", THREADS, "| batch:", BATCH, "\n");
let sink = 0; boxplot(() => { summary(() => { bench(`host (${REQUESTS.toLocaleString()} req)`, () => { sink = runHost(batches).budgetedTokens; });
bench( `knitting (${THREADS} threads, ${REQUESTS.toLocaleString()} req)`, async () => { sink = (await runWorkers(callBatch, batches)).budgetedTokens; }, ); }); });
await run(); console.log("last budgeted tokens:", sink.toLocaleString());}
function runHost(batches: PromptInput[][]): PromptBudgetSummary { return mergeSummaries(batches.map(summarizeBatchHost));}
async function runWorkers( callBatch: (inputs: PromptInput[]) => Promise<PromptBudgetSummary>, batches: PromptInput[][],): Promise<PromptBudgetSummary> { return mergeSummaries(await Promise.all(batches.map(callBatch)));}
if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; });}When this matters
Section titled “When this matters”Token budgeting is a preflight step that runs on every LLM request, and it is pure CPU work on the thread you most want free. On a busy chat service — many users, long histories — it is enough to add latency to requests that are otherwise just waiting on the model. Moving it to a pool keeps the main thread on routing and I/O, and gives you predictable input sizes, which is what makes cost per request predictable too.