protologue

Many-shot Jailbreaking

Many-shot jailbreaking fills a long context window with many fabricated dialogue examples in which an assistant complies with harmful requests, exploiting in-context learning to override the model's safety training.

Description

Anthropic researchers found the attack's effectiveness followed a power law in the number of shots, mirroring the scaling of benign in-context learning.

Sources

  1. Anthropic (2024). Many-shot jailbreaking.

Cite this entry

Protologue. (2026). Many-shot Jailbreaking. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0092). https://protologue.com/t/many-shot-jailbreaking/

BibTeX
@misc{protologue_many_shot_jailbreaking,
  title = {Many-shot Jailbreaking},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0092},
  url = {https://protologue.com/t/many-shot-jailbreaking/}
}

Markdown JSON