Prompt Compression
Also called LLMLingua, context compression.
Prompt compression shortens a prompt by removing tokens that contribute little information, typically scored by a smaller language model, to cut cost and latency while preserving task performance.
Description
LLMLingua used a small model's perplexity to drop low-information tokens and reported high compression ratios with limited performance loss.
Sources
- Jiang et al. (2023). LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.
Cite this entry
Protologue. (2026). Prompt Compression. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0080). https://protologue.com/t/prompt-compression/
BibTeX
@misc{protologue_prompt_compression,
title = {Prompt Compression},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0080},
url = {https://protologue.com/t/prompt-compression/}
}