A classification API should answer a precise question and make the limits of that answer visible. Returning a label is easy. Defining the label, interpreting a score, and deciding whether the application should act on the result require more care. Those decisions matter whether classification uses a dedicated model, an embedding pipeline, or a prompted language model.
Consider an illustrative support-routing task with the labels “billing,” “technical,” and “general.” The API's purpose is to help route messages, while uncertain cases remain reviewable. This article develops that example from its label contract through evaluation. The classification Tensor API guide gives the surrounding workflow.
Define the decision before choosing the model
Write one sentence describing what each label includes and excludes. A billing complaint that also mentions an application error will otherwise receive inconsistent labels from both annotators and models. Decide whether one label must win, whether multiple labels are allowed, or whether the system should abstain when the available context is insufficient.
Single-label multiclass classification chooses one class from a set. Multilabel classification allows several classes to apply at once. These tasks need different output and evaluation rules. A list of independent label scores should not be interpreted as a distribution over mutually exclusive alternatives merely because the same numerical range appears in both interfaces.
Keep an “insufficient information” outcome distinct from a miscellaneous content category. “General” can describe a real kind of request; “needs review” describes the system's decision process. Conflating the two hides uncertainty inside a label and makes it hard to measure how much traffic the automation actually handles.
Keep the label map beside the tensor
A classifier may produce a score tensor shaped [batch, classes]. The column order needs a versioned label map. For a token classifier, an additional sequence axis may be present, which changes the meaning of both the output and any aggregation. Inspect the model's actual task definition before converting its tensors into a public response.
Preserve input identifiers when batching. The response for message seventeen must remain associated with message seventeen even if internal scheduling reorders work. Include the label-schema revision and model revision in the result or associated audit record. A changed definition of “technical” can affect routing even when the model architecture and JSON fields stay the same.
For a prompted LLM, constrain the returned labels and validate them before accepting the result. A natural-language explanation can help a reviewer, but it should not silently introduce a fourth label or override application routing rules. The AI LLM Tensor API workflow separates generation from these downstream decisions.
Do not confuse a model score with a reliable probability
A classifier score can be useful for ordering examples without accurately describing the chance that a label is correct. Applying softmax to logits creates a normalized distribution, but normalization alone does not establish calibration. A language model's written statement that it is “very confident” is also not a measured probability.
The scikit-learn probability-calibration guide explains calibration in terms of agreement between predicted probabilities and observed outcomes. Among a sufficiently representative group receiving similar probability estimates, the observed fraction of positives should be close to those estimates. Reliability diagrams examine that relationship across probability ranges.
Assess calibration using examples separate from those used to fit the underlying model and tune the final decision rule. A calibration method can improve a score's interpretation on the evaluated distribution without guaranteeing the same behavior on future inputs. Report the population and period used for the assessment, especially when language, customer mix, or document sources can change.
If a score has not been validated as a probability, name it accordingly. A field such as model_score with a documented interpretation is more useful than an unexplained “confidence” percentage. Consumers need to know whether comparisons are meaningful across labels, models, or versions before building thresholds around that field.
Choose thresholds from the cost of mistakes
A routing threshold is an application policy. Raising it can send more cases to review; lowering it can automate more cases while changing the error profile. Measure that tradeoff against the task instead of borrowing a threshold from another project. Different labels may need different policies if their mistakes have different consequences.
For example, suppose a billing queue has limited reviewer capacity. Evaluate both the rate of wrongly routed messages and the number of messages deferred for review. A system with a strong score on automatically accepted examples may still be impractical if it defers almost everything. Coverage, the fraction handled automatically, belongs beside accuracy among accepted decisions.
Define what happens when review capacity is exhausted. Queueing, returning an explicit pending state, or using a documented fallback are design choices. Silently treating uncertain cases as confident predictions changes the policy precisely when the system is under pressure.
Build an evaluation set that resembles future traffic
Collect examples that represent the intended languages, input lengths, sources, and ambiguous cases. Write annotation guidance and resolve disagreements before treating labels as ground truth. If trained reviewers cannot apply a category consistently, refining the category definitions may improve the system more than switching models.
Avoid leaking closely related records between training and evaluation. Multiple messages from one ticket or near-duplicate documents can make a random split look easier than real deployment. Use groups when related examples should remain together. For changing traffic, a later time period can provide a more realistic test of performance than a purely shuffled split.
Reserve a final evaluation set that does not guide prompt edits, threshold selection, or calibration fitting. Repeatedly inspecting the same test failures and adjusting the system turns that set into development data. Keep a compact development set for iteration, and maintain a separate check for the final candidate.
Keep a slice report for cases that are easy to overlook: short messages with little context, messages containing two requests, and inputs from a source absent during development. Report small sample sizes instead of treating an empty error count as proof of reliability. When one slice performs poorly, investigate its labeling and preprocessing before applying a global threshold change that could affect every other group.
Read several metrics together
Precision asks how many predicted positives were correct. Recall asks how many actual positives were found. A confusion matrix shows which classes are being confused. Overall accuracy can hide a rare class that the system almost never identifies, so inspect results per label and report the number of examples behind each result.
| Illustrative billing evaluation | Count |
|---|---|
| Correctly predicted billing | 36 |
| Incorrectly predicted billing | 4 |
| Billing messages missed | 9 |
These fictional counts illustrate the arithmetic, not measured model performance. Billing precision is 36 / (36 + 4) = 90%. Billing recall is 36 / (36 + 9) = 80%. The two numbers answer different operational questions. Report support and uncertainty when a small sample makes a percentage unstable.
For summaries across classes, macro averaging gives classes equal weight, while weighted averaging accounts for their support. State the averaging method. Evaluate abstentions and unavailable responses separately so that dropping difficult cases does not create an artificially flattering quality score.
Make the response operationally clear
The following illustrative object separates a raw prediction from the decision about what happens next. Its score is deliberately named as a model score rather than a calibrated probability.
{
"item_id": "example-17",
"predicted_label": "billing",
"model_score": 0.73,
"score_kind": "uncalibrated",
"decision": "needs_review",
"label_schema": "support-routing-v1",
"model_revision": "local-demo-v1"
}
Use a separate error state when classification was not completed. A timeout is not evidence that an item belongs in the general category. For generated responses, the structured-output validation guide explains how to reject invalid fields and incomplete objects before they enter a workflow.
Conclusion: evaluate the whole decision system
A dependable classification interface joins explicit labels, interpretable scores, measured thresholds, and honest failure states. Track performance after deployment using reviewed samples and compare it across relevant slices of traffic. Revisit the label contract when the task changes, and reevaluate whenever a prompt, model, or preprocessing revision changes behavior. The useful outcome is a routing decision whose quality and coverage can be explained.


