protologue

Token

Also called subword, BPE token.

A token is the basic unit of text a language model reads and writes, typically a word, word fragment, or character sequence produced by a subword tokenizer such as byte-pair encoding.

Description

Context limits, pricing, and generation speed are all measured in tokens. Because tokenization splits text unevenly, character-level tasks such as counting letters or reversing strings are harder for models than they appear, and the same content can cost different numbers of tokens in different languages.

Sources

  1. Sennrich et al. (2015). Neural Machine Translation of Rare Words with Subword Units.

Cite this entry

Protologue. (2026). Token. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0005). https://protologue.com/t/token/

BibTeX
@misc{protologue_token,
  title = {Token},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0005},
  url = {https://protologue.com/t/token/}
}

Markdown JSON