Follow the thread

Token budgets.

Make context and usage decisions with an explicit accounting model.

This tag brings together inference and token-budget guides for teams deciding what to send, what to reserve for a response and what to measure afterward. The articles avoid fixed assumptions about how many words fit in a token. The relevant unit depends on the model and its tokenizer.

Begin by distinguishing an estimate made before a request from usage returned afterward. Keep the request configuration and relevant provider semantics with both. That record makes it possible to compare prompt changes without silently changing the task. Budget decisions should also preserve enough evidence and output capacity for a useful answer.

2 articles tagged Token budgets

Articles tagged Token budgets

Explore related subjects