WHAT YOU WILL BUILD
A context selection that fits a supplied budget while keeping the current task and constraints.
A context window is a capacity limit, not a target to fill. Budget the fixed instructions, tool definitions, current task, evidence, history, and expected output. The exercise uses supplied counts so you can inspect the selection arithmetic without a tokenizer.
Before you start
Use a text editor and Node.js 22 or newer ↗. Check your version with node --version. Save the downloaded file in an empty folder, then open a terminal in that folder.
All data is included. No packages, account, or API key are required. Token counts and the 1,000-token window are invented teaching inputs, not a provider specification or token estimate for the displayed text.
Download the lab (.mjs)FOLLOW ALONG
Work through the example.
- 01
Reserve fixed costs
Start with a 1,000-token allowance, reserve 200 for output and 150 for fixed input. Calculate the remaining 650 before choosing history or evidence.
- 02
Select by task importance
The sample order is constraints, current task, evidence, then old chatter. Run the script to retain the first three entries, totaling 550 tokens, and leave 100 unallocated.
- 03
Test an impossible budget
Increase constraints to 800. Observe that this simple selector skips them. A real system must fail or request a shorter task when required material cannot fit; it must not silently discard required constraints.
- 04
Measure a real request later
Use the selected provider's tokenizer or counting endpoint with the exact messages and tools. Recheck the total after formatting; include any provider-specific accounting for images or reasoning before making cost claims.
THE COMPLETE LAB
Run it locally.
Run this command from the folder containing your downloaded file:
node prompt-context-budget.mjsView or copy the complete JavaScript
// SaveMyToken local lab. Run with Node.js 22 or newer.
import assert from "node:assert/strict";
// Supplied token counts, not a tokenizer or measured provider usage.
const windowSize = 1000, outputReserve = 200, fixedInput = 150;
const available = windowSize - outputReserve - fixedInput;
const items = [
{ id: "constraints", tokens: 120 },
{ id: "current-task", tokens: 180 },
{ id: "evidence", tokens: 250 },
{ id: "old-chatter", tokens: 400 },
];
let used = 0;
const kept = items.filter(item => {
if (used + item.tokens > available) return false;
used += item.tokens;
return true;
});
assert.ok(used + fixedInput + outputReserve <= windowSize);
assert.equal(used, 550);
console.log("Available input budget:", available);
console.log("Kept:", kept.map(item => item.id).join(", "));
console.log("Unallocated tokens:", available - used);
Expected output for the unchanged example
Available input budget: 650
Kept: constraints, current-task, evidence
Unallocated tokens: 100The assertions also check the baseline behavior. After changing an input, predict the result and update the relevant assertion.
NOW CHANGE ONE THING
Make the example your own.
Mark constraints and current-task as required. Add a check that throws when required items exceed the available budget, and verify it with the 800-token case.
You are done when…
The baseline fits, and your extended version explicitly rejects an impossible required-input budget instead of dropping instructions.
If something goes wrong
Characters, words, and tokens are different units. This script never counts the text itself. Keep required tool-call/result pairs together when adapting selection to a real conversation.
EXPLAIN WHAT YOU LEARNED
Interview practice
Try answering aloud before opening the reference answer. These are original learning questions, not a record of any employer's interviews.
01What belongs in a context budget?
Include fixed instructions, tool definitions, messages, retrieved evidence, and space for output. Verify model-specific handling of other inputs and reasoning with the provider.
Watch for: Counting only the most recent user message understates the request.
02When is summarization useful?
When older material contains useful decisions but too much detail to resend. Preserve constraints, unresolved issues, and retrievable source references, then check the summary for omissions.
Watch for: Summarization has a cost and can remove information needed later.
03What should remain exact during compression?
Keep identifiers, user constraints, critical numbers, code that must run, and evidence whose wording matters. Compress surrounding explanation instead.
Watch for: An approximate paraphrase can change a permission or a requirement.
04How is prompt caching different from history trimming?
Caching may reuse processing of eligible repeated input; trimming changes which input is sent. Their billing and quality effects differ and must be measured separately.
Watch for: A cached request can still include a large context and paid output.
05What if the necessary evidence does not fit?
Narrow the task, retrieve a better subset, process it in stages, or choose a suitable model after checking its limits. Make missing evidence explicit.
Watch for: Silently dropping the exception to a rule may produce a confident wrong answer.
Sources and further reading
The explanation and local exercises were written for SaveMyToken. These references support the underlying concepts; the sample outputs describe only the supplied examples.
Anthropic: context windows ↗LangChain: short-term memory ↗