
Daniel Okonkwo
Cost is a property of the architecture rather than of the prompt, and it can be read off a design before a single request is sent. Output is typically several times the price of input, which means a system generating long answers costs more than the same system summarising, and a summarisation step that runs on every document can dominate the bill. Caching and batch pricing both work by lowering the effective rate for predictable work. Cached input suits a system prompt that never changes. Batch suits anything that can tolerate delay. Between them, they are the cheapest lever available and they require no model change at all. Embeddings are priced per token too, and they are the cost people forget, because the indexing step runs once per document while the querying step runs on every request. A corpus that is cheap to embed and expensive to search is a common shape, and it changes which retrieval strategy makes sense. Images have their own unit, priced per generated image and varying with resolution. Any estimate has to state which it assumed. The number to be careful with is the token estimate, because tokenizers differ across model families and a character count is not a token count. I show the range an estimate carries, not a single figure. I add a worked example end to end, from a request with a stated token count to a monthly figure, because the arithmetic is the point and nobody should have to reconstruct it from three numbers on a pricing page.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →