
Kenji Morita
I spend my time on the arithmetic behind a model call, because the headline numbers and the numbers that matter are different numbers. A context window is the total of everything the model sees in one call: the system prompt, the conversation so far, any documents retrieved for this request, the tool or function definitions, and the output you have reserved. Vendors quote the total. What you have available is that figure minus everything else, and on a tool-using application the tool definitions alone can take a sizeable slice. Output reservation is the part people forget. A model given a large context and a large output allowance may still be told to reserve the output budget, so the usable input is smaller than the window suggests. Asking for the maximum output on a long conversation is how you get a truncated answer. Tokenization varies by model family, so an estimate carried over from one model to another is an estimate of the wrong quantity. Code and structured output pack differently from prose, and a JSON document full of short keys can tokenize far less efficiently than its character count suggests. Temperature and sampling parameters are worth covering too, because they interact with everything else. A request for a deterministic output needs the lowest setting available, and a tool-calling loop wants that too, since a sampled argument is a failed call. I finish on stop sequences, which quietly truncate output when they appear inside the content the model is producing. They are a blunt instrument and belong only where you control the output format.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →