Follow the thread

Evaluation.

Define useful behavior before choosing a success metric.

Evaluation connects a model result to the task it is supposed to perform. This tag covers classification decisions, structured-output validation and the evidence needed for tokenized-data workflows. Start with the outcome that matters to a caller, then work backward to the examples and checks that make it testable.

Keep representative cases separate from examples used to tune a prompt or policy. Include missing fields, ambiguous inputs and unusual but valid records. The guides distinguish structural validity from semantic correctness: an output can satisfy a schema and still fail the task. A review path is part of the design when uncertainty cannot be resolved automatically.

3 articles tagged Evaluation

Articles tagged Evaluation

Explore related subjects