A token budget is a rule for deciding how much work an LLM request may consume before the application sends it. It connects context selection, output length, concurrency, retries, and operational cost. Without that rule, one unusually long document or a loop of correction attempts can turn a modest task into an unpredictable workload.
Batching and caching can help, but they solve different problems. Batching organizes work; caching reuses eligible computation or results. Neither replaces a clear request contract. The tokens Tensor API guide explains the underlying vocabulary. Here, we develop an illustrative budgeting policy that you can adapt to your own application.
Separate the limits you are trying to satisfy
A request can be limited by the model's context capacity, its permitted output length, a provider's rate limits, your own concurrency policy, and an application spending allowance. These constraints are related but not interchangeable. A request that fits the context can still wait for capacity or exceed the amount of work allowed for one task.
Document how the selected model and provider count context, including any special handling for tools, images, or reasoning features that you use. Do not assume that a visible text length covers the entire request. Review those rules when changing model versions or integration paths.
For a hypothetical model with a shared input-and-output context limit, a planning rule might be:
assembled_input_tokens
+ reserved_output_tokens
+ safety_margin
<= documented_context_capacity
The margin handles estimation uncertainty and small assembly changes; it is not extra model capacity. For an illustrative capacity of 12,000 tokens, reserving 2,000 for output and 1,000 as margin leaves 9,000 for assembled input. Those figures are teaching values, not the specification of a named model.
Estimate early and reconcile with actual usage
Use a compatible tokenizer or a provider-supported counting operation when available. Count the assembled request close to submission, after instructions, retrieved content, conversation history, and tool definitions have been selected. Counting only the user's visible question can omit most of the actual input.
Anthropic's token-counting documentation describes its count as an estimate that can differ slightly from actual input usage. That distinction is useful beyond any one provider: keep estimated input tokens in a separate field from the usage later reported for the completed request. An estimate supports admission decisions; a usage record describes what happened.
A rough character-based estimate can help an early interface decide whether a document is obviously too large, but it should not become the authoritative counter. Test text from the languages and formats you expect. Source code, tables, and mixed-language documents may behave differently from ordinary English prose.
Reduce context according to relevance
When input exceeds the budget, apply an explicit policy instead of removing an arbitrary suffix. Preserve instructions that define the task and identify the evidence needed to answer it. Select relevant passages, remove duplicates, and shorten material that contributes little to the requested result.
A summary can reduce context size, but it introduces a new transformation to evaluate. Store references back to the original material and test whether important qualifications survive. If a required fact disappears during compression, the shorter prompt may be cheaper while producing a worse answer.
Make the reduction observable. Record how many passages were selected, whether truncation occurred, and which preprocessing revision was used. A debugging session should be able to distinguish “the model ignored the evidence” from “the application never supplied the evidence.” The LLM inference pipeline guide follows these input transformations in more detail.
Batch for the workload's actual deadline
Interactive requests need a different scheduling policy from a collection of reports due later. In your own model runtime, grouping compatible sequences can reduce per-example overhead, while mismatched lengths can increase padding and memory use. Group by model and processing contract before considering sequence length.
A provider's asynchronous batch interface is a different mechanism from building one runtime tensor batch. Check its documented completion, cancellation, and result-retrieval behavior. Keep a stable item identifier so that outputs can be matched to inputs even when completion order differs.
Set bounds on queue waiting and batch size. Measure whether a larger group improves useful throughput enough to justify added delay. If one oversized item causes a whole group to fail, isolate or reject it according to policy. A batch should have explicit per-item success and failure handling rather than one vague status that loses partial results.
Cache the thing you actually intend to reuse
Distinguish three mechanisms: runtime state reused during generation, provider prompt caching, and an application cache of completed responses. They have different keys, validity rules, and consequences. A completed-answer cache must consider whether the requested information or the user's permissions have changed.
The official Claude prompt-caching guide describes reuse of matching prompt prefixes and requires exact matching of the relevant segments. Put stable, reusable material before request-specific content when that arrangement also fits the task. Changing a timestamp inside the eligible prefix can prevent matching.
Keep the full logical input in the context budget even when part of its processing is cached. Prompt caching does not mean the source disappeared from the request or that generated output can be skipped. Check documented model eligibility, cache lifetimes, minimum sizes, and usage fields instead of assuming that every repeated request qualifies.
For an application response cache, include the prompt revision, model identity, source revision, and relevant authorization scope in the key design. Define expiration and invalidation behavior. Reusing a stale summary can be incorrect even when the text of the latest question is identical.
Budget attempts as part of one logical task
A task may involve the initial generation, one validation correction, and a retry after a transient failure. Give the whole task a ceiling as well as each individual request. Otherwise every attempt can satisfy its local limit while the sequence exceeds the intended allowance.
Record why an attempt occurred. A rate-limit retry, a schema correction, and a model fallback represent different failure modes. Use a bounded retry policy with the provider's documented guidance, and avoid immediately resubmitting requests that failed because their input was invalid.
Cancellation and network timeouts require careful accounting. The client losing a connection does not prove that no processing occurred. Keep uncertain outcomes visible until provider records or application reconciliation resolve them. Do not automatically classify every timeout as zero usage.
Keep usage records that answer practical questions
A useful internal record joins the logical task, individual attempt, model, prompt revision, estimated input size, reported usage, completion state, and duration. Preserve provider-specific usage detail where necessary, then derive clearly named application totals. Do not sum fields that overlap simply because they all contain token counts.
| Record field | Purpose |
|---|---|
| Task and attempt identifiers | Connect retries to the original request. |
| Estimated input tokens | Explain the admission decision. |
| Reported usage categories | Reconcile actual processing. |
| Cache and batch status | Explain which optimizations applied. |
| Terminal state and duration | Separate useful completion from failure. |
If you calculate spending, use the applicable provider rates and their effective dates, and retain the accounting assumptions. A token total alone does not identify a billable amount. The AI token usage guide covers how to keep consumption concepts clear across integrations.
Optimize against useful completion
Compare policies using representative tasks and a stable quality check. Measure completed useful results, correction rate, queue delay, total task duration, and consumption. A smaller output limit can reduce generated tokens while increasing incomplete answers and retries. A cache can improve repeated workloads while offering little benefit to unique prompts.
Review the distribution as well as the average. A few long-running tasks can dominate the experience of a subset of users. Tune limits for recognizable workload classes instead of hiding very different jobs behind one global default.
Conclusion: make consumption explainable
Start with an explicit task budget, count the assembled input, and reserve room for the required output. Apply batching and caching only where their behavior matches the workload. Keep estimates, attempts, and reported usage connected so that a cost or latency change has an explanation. For provider-specific implementation details, continue with the Anthropic API integration guide.


