Prompt Caching
Also called context caching, prefix caching.
Prompt caching stores the model's processed state for a reused prompt prefix, such as a long system prompt or document, so later requests sharing that prefix are cheaper and faster.
Description
Because caches match on an exact prefix, prompts are structured with stable content first and variable content last.
Sources
- Anthropic (2024). Prompt caching.
Cite this entry
Protologue. (2026). Prompt Caching. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0081). https://protologue.com/t/prompt-caching/
BibTeX
@misc{protologue_prompt_caching,
title = {Prompt Caching},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0081},
url = {https://protologue.com/t/prompt-caching/}
}