Follow the thread

Tokenization.

Understand the conversion from text to model-specific input units.

Tokenization belongs to the representation layer of a language-model workflow. This collection separates readable text, token IDs, embeddings and the tensors that carry numerical values. Read the concept guide when discussions use these terms as if they meant the same thing.

For a concrete example, take a short input and track the identity of the tokenizer and the model that will consume its result. Note where special tokens, padding or truncation could affect the representation. The purpose is to document the transformation, not to assume a universal correspondence between a word, a token and a vector. Related inference guides show how those distinctions matter downstream.

1 article tagged Tokenization

Articles tagged Tokenization

Explore related subjects