From inference to applications
Connect LLM applications to the right outputs
An LLM application usually needs an answer, a classification, or a structured record. The underlying model works with token identifiers, hidden representations, and prediction scores. An API contract connects those layers, and it does not need to expose every internal value to be useful. This guide helps you decide what your own integration should accept and return. Start with the consumer’s task, check the selected runtime or provider’s actual capabilities, and preserve the metadata needed to understand how each result was produced.
Start from the application result
Choose the output that the consumer can use directly: generated text, an approved label, a validated object, or an embedding. Each choice has a different acceptance rule. Text may require factual review; a label must belong to a vocabulary; a structured object needs schema and semantic validation.
Keep application-owned fields outside generated content. Request identifiers, selected model revisions, and measured durations should come from the surrounding system. The model does not need to invent information that the application already knows.
- Name the task and required input context.
- Define success, unavailable information, and failure.
- Specify the output type and its validation rule.
- Record the revisions that can change behavior.
This compact contract helps callers integrate without depending on how the model arranges its intermediate representations.
Understand the internal representations
Token identifiers commonly form a tensor with batch and sequence axes. Hidden representations add a hidden-width axis. A causal language-model output can contain a vocabulary score for each represented position. These layouts explain the computation, but they are not interchangeable application results.
An embedding is a numerical representation chosen for a downstream task. A vocabulary score supports a decoding decision. A generated string is the result of decoding a continuation. Calling all three a “tensor response” hides distinctions that consumers need.
Model implementations may expose different optional fields. Check the specific class or endpoint before assuming access to logits, hidden states, or attention values. Request additional outputs only when a downstream workflow has a clear use for them, and document their dimensions and model dependence when you do expose them.
Return a result that can be diagnosed
Separate a useful completion from a successful network exchange. A response may be truncated, cancelled, unavailable, or invalid for the application even when transport succeeds. Preserve the provider or runtime status, then map it into explicit application states.
{
"item_id": "example-42",
"state": "completed",
"output_kind": "validated_object",
"prompt_revision": "document-summary-v1"
}This illustrative envelope leaves the actual content and usage fields to the selected contract. Keep estimated input size distinct from reported usage, and connect any correction attempts to the original task. Retain sufficient metadata to compare model or prompt changes on a stable evaluation set.
When streaming, define how consumers distinguish an unfinished fragment from a final validated result. Delay actions that require a complete object until the terminal state and validation checks are available.
Use the official reference for the documented interface; the workflow recommendations above are Tensor API Lab guidance.