Skip to content

LLM input shaping

This example trims LLM prompts to fit a token budget before an API call. It counts tokens with tiktoken, drops the oldest conversation turns first, and shortens the query only if removing every turn is still not enough. The budgeting logic is model-agnostic — the example uses the tokenizer for gpt-4o-mini, but the pattern is the same for any model with an input limit.

  1. The host builds synthetic chat traffic: same system prefix, 3 to 14 turns of history, and a query.
  2. Each task builds the full prompt and counts tokens with tiktoken.
  3. If the prompt is over budget, it drops the oldest turns one at a time, re-counting after each.
  4. If no turns are left and it is still over, it clips the query to whatever budget remains.
  5. The host aggregates token savings and trim counts.

Four files:

  • token_budget.ts — the budgeting logic and the two tasks
  • prompt_fixtures.ts — the synthetic conversations, so both scripts share one generator
  • run_prompt_token_budget.ts — what budgeting actually does to a workload
  • bench_prompt_token_budget.ts — host vs workers, with mitata

There are two tasks because the right return shape depends on what you need. preparePrompt returns the whole PromptPlan, prompt string included, which is what an application wants. summarizeBatch takes a batch and returns counters only — no prompt strings — which is what you want when you are measuring, or when you only need the accounting.

Minimal usage:

using pool = createPool({ threads: 3 })({ preparePrompt });
const plan = await pool.call.preparePrompt({
model: "gpt-4o-mini",
systemPrefix: "You are a docs assistant.",
history: ["Need guidance on schema validation.", "Keep the answer short."],
query: "Give a migration plan and one code example.",
maxInputTokens: 400,
});

Output shape:

{
prompt: "...final prompt string...",
rawInputTokens: 1061,
inputTokens: 399,
staticTokens: 47,
dynamicTokens: 352,
trimmedTurns: 3,
queryWasTrimmed: true,
}

The useful part is not just the final prompt. You also get the bookkeeping needed to explain why a prompt was trimmed and by how much — which turns were dropped, whether the query itself had to be clipped, and how much of the budget the fixed prefix uses.

bun.sh
bun src/run_prompt_token_budget.ts

Expected output:

Prompt token budgeting
requests : 2,000
budget : 400 tokens per prompt
raw tokens : 904,702
budgeted tokens : 684,035
saved : 220,667 (24.4%)
trimmed : 1,057 / 2,000 requests
turns dropped : 4,286
query clipped : 118
elapsed : 1640 ms
One request, in detail:
1061 -> 399 tokens
3 turns dropped
query clipped: true
47 of those tokens are the fixed prefix

Choose a budget that some sample prompts exceed. Otherwise, the example only measures token counting and never exercises the trimming logic.

bun.sh
bun src/bench_prompt_token_budget.ts

The benchmark compares budgeting throughput on the host with the same work on workers. Batches are sent through summarizeBatch, so only counters come back, never prompt strings.

token parity check: host=341,835 worker=341,835 OK match
Prompt token budgeting benchmark (mitata)
workload: build prompt + tokenize + trim to budget
requests per iteration: 1,000
budget: 400 tokens
trimmed: 528 requests, 24.4% of tokens saved
threads: 3 | batch: 32
benchmark avg (min … max)
host (1,000 req) 1.45 s/iter (1.39 s … 1.46 s)
knitting (3 threads, 1,000 req) 471.93 ms/iter (468.14 ms … 497.00 ms)
summary
knitting (3 threads, 1,000 req)
3.08x faster than host (1,000 req)

This workload stops getting faster after a few workers. The measurements below show why choosing a pool size requires more than counting cores.

On a 4-core laptop the best setting measured here was three workers plus the inliner — four lanes for four cores — at roughly 3x. Beyond that, performance drops: eight workers measured consistently slower than three or four. One worker performed about the same as the host, within measurement noise.

With plain arithmetic in place of tokenization, the same pool and batching reached 4.8x at eight threads on the same laptop. This suggests that the tokenizer is limiting scaling in this workload.

The reason is that tiktoken is a WASM module with a large BPE table, and every worker instantiates its own copy. Resident memory here went from 67 MB at baseline to 630 MB with eight workers — roughly 55-60 MB per worker. Those tables are not shared and exceed the available CPU cache. Adding workers can therefore increase cache pressure without improving throughput.

Two practical consequences:

  • Size the pool for the tokenizer, not for os.cpus().length. Start at 2-4 lanes and measure. The inliner helped in this benchmark — it adds a lane without adding another encoder-sized worker.
  • Keep many more batches than lanes. Batching amortizes dispatch, but one giant batch per lane cannot rebalance. At 1,000 requests in batches of 250 the pool measured slower than the host (0.81x); the same work in batches of 32 was the fastest setting tried.
token_budget.ts
import { task } from "knitting";
import { encoding_for_model } from "tiktoken";
export type PromptInput = {
model: string;
systemPrefix: string;
history: string[];
query: string;
maxInputTokens: number;
};
export type PromptPlan = {
prompt: string;
rawInputTokens: number;
inputTokens: number;
staticTokens: number;
dynamicTokens: number;
trimmedTurns: number;
queryWasTrimmed: boolean;
};
/** What a batch of plans adds up to. No prompt strings, so it is cheap to return. */
export type PromptBudgetSummary = {
rawTokens: number;
budgetedTokens: number;
staticTokens: number;
dynamicTokens: number;
trimmedRuns: number;
queryTrimmedRuns: number;
turnsDropped: number;
};
type Encoder = ReturnType<typeof encoding_for_model>;
const decoder = new TextDecoder();
const encoderCache = new Map<string, Encoder>();
const staticTokenCache = new Map<string, number>();
// A tiktoken encoder is a WASM instance holding its own BPE table, and it costs
// tens of megabytes resident. This cache lives per worker, so the real ceiling
// is MAX_ENCODERS * threads. Keep it small, and free whatever you evict.
const MAX_ENCODERS = 2;
function getEncoder(model: string): Encoder {
const cached = encoderCache.get(model);
if (cached) return cached;
const enc = encoding_for_model(model as never);
encoderCache.set(model, enc);
if (encoderCache.size > MAX_ENCODERS) {
const oldest = encoderCache.keys().next().value!;
encoderCache.get(oldest)!.free();
encoderCache.delete(oldest);
}
return enc;
}
/** The system prefix is identical on every request, so tokenize it once. */
function getStaticTokens(model: string, prefix: string, enc: Encoder): number {
const key = `${model}\x1f${prefix}`;
const cached = staticTokenCache.get(key);
if (cached !== undefined) return cached;
const value = enc.encode(prefix).length;
staticTokenCache.set(key, value);
return value;
}
export function clearPromptBudgetCaches(): void {
for (const enc of encoderCache.values()) enc.free();
encoderCache.clear();
staticTokenCache.clear();
}
function normalizeText(value: string): string {
return value.replace(/\s+/g, " ").trim();
}
function buildPrompt(prefix: string, history: string[], query: string): string {
const rows = [prefix.trim(), "", "Conversation context:"];
for (let i = 0; i < history.length; i++) {
rows.push(`- Turn ${i + 1}: ${history[i]}`);
}
rows.push("", `User request: ${query}`);
return rows.join("\n");
}
function truncateToTokenBudget(
enc: Encoder,
text: string,
maxTokens: number,
): string {
if (maxTokens <= 0) return "";
const tokens = enc.encode(text);
if (tokens.length <= maxTokens) return text;
return decoder.decode(enc.decode(tokens.slice(0, maxTokens)));
}
/**
* Fit a prompt inside `maxInputTokens`: drop the oldest turns first, and only
* clip the query if dropping every turn still is not enough.
*/
export function preparePromptHost(input: PromptInput): PromptPlan {
const maxInputTokens = Math.max(64, input.maxInputTokens);
const history = input.history.map(normalizeText).filter(Boolean);
const enc = getEncoder(input.model);
const staticTokens = getStaticTokens(input.model, input.systemPrefix, enc);
let query = normalizeText(input.query);
let prompt = buildPrompt(input.systemPrefix, history, query);
const rawInputTokens = enc.encode(prompt).length;
let inputTokens = rawInputTokens;
let trimmedTurns = 0;
let queryWasTrimmed = false;
while (inputTokens > maxInputTokens && history.length > 0) {
history.shift();
trimmedTurns++;
prompt = buildPrompt(input.systemPrefix, history, query);
inputTokens = enc.encode(prompt).length;
}
// No turns left to drop and still over: the query itself is the problem.
if (inputTokens > maxInputTokens) {
const scaffolding = buildPrompt(input.systemPrefix, history, "");
const remaining = Math.max(
16,
maxInputTokens - enc.encode(scaffolding).length,
);
const clipped = truncateToTokenBudget(enc, query, remaining);
queryWasTrimmed = clipped.length < query.length;
query = clipped;
prompt = buildPrompt(input.systemPrefix, history, query);
inputTokens = enc.encode(prompt).length;
}
return {
prompt,
rawInputTokens,
inputTokens,
staticTokens,
dynamicTokens: Math.max(0, inputTokens - staticTokens),
trimmedTurns,
queryWasTrimmed,
};
}
/**
* Budget a whole batch and return only the counters. Batching amortizes the
* per-call dispatch, and dropping the prompt strings keeps the return small.
*/
export function summarizeBatchHost(
inputs: PromptInput[],
): PromptBudgetSummary {
const totals = emptySummary();
for (let i = 0; i < inputs.length; i++) {
const plan = preparePromptHost(inputs[i]!);
totals.rawTokens += plan.rawInputTokens;
totals.budgetedTokens += plan.inputTokens;
totals.staticTokens += plan.staticTokens;
totals.dynamicTokens += plan.dynamicTokens;
totals.turnsDropped += plan.trimmedTurns;
if (plan.trimmedTurns > 0) totals.trimmedRuns++;
if (plan.queryWasTrimmed) totals.queryTrimmedRuns++;
}
return totals;
}
export function emptySummary(): PromptBudgetSummary {
return {
rawTokens: 0,
budgetedTokens: 0,
staticTokens: 0,
dynamicTokens: 0,
trimmedRuns: 0,
queryTrimmedRuns: 0,
turnsDropped: 0,
};
}
export function mergeSummaries(
parts: PromptBudgetSummary[],
): PromptBudgetSummary {
return parts.reduce((a, b) => ({
rawTokens: a.rawTokens + b.rawTokens,
budgetedTokens: a.budgetedTokens + b.budgetedTokens,
staticTokens: a.staticTokens + b.staticTokens,
dynamicTokens: a.dynamicTokens + b.dynamicTokens,
trimmedRuns: a.trimmedRuns + b.trimmedRuns,
queryTrimmedRuns: a.queryTrimmedRuns + b.queryTrimmedRuns,
turnsDropped: a.turnsDropped + b.turnsDropped,
}), emptySummary());
}
/** Returns the full plan, prompt string included. */
export const preparePrompt = task<PromptInput, PromptPlan>({
f: preparePromptHost,
});
/** Returns counters only. This is the one worth benchmarking. */
export const summarizeBatch = task<PromptInput[], PromptBudgetSummary>({
f: summarizeBatchHost,
});

Token budgeting is a preflight step that runs on every LLM request, and it is pure CPU work on the thread you most want free. On a busy chat service — many users, long histories — it is enough to add latency to requests that are otherwise just waiting on the model. Moving it to a pool keeps the main thread on routing and I/O, and gives you predictable input sizes, which is what makes cost per request predictable too.