Prompt Optimization
Automatic and learned methods that search for, compress, or train better prompts.
- Automatic Prompt EngineerPTL-0073
- Directional Stimulus PromptingPTL-0079
- DSPyPTL-0075
- Emotion PromptingPTL-0082
- Low-Rank AdaptationPTL-0078
- Optimization by PromptingPTL-0074
- Prefix TuningPTL-0077
- Prompt CachingPTL-0081
- Prompt CompressionPTL-0080
- Prompt TuningPTL-0076
Definitions
- Automatic Prompt Engineer
- Automatic Prompt Engineer (APE) uses a language model to generate candidate instructions for a task from input-output examples, scores each candidate on held-out data, and selects the best one.
- Directional Stimulus Prompting
- Directional stimulus prompting trains a small policy model to generate instance-specific hints, such as keywords, that are added to the prompt to steer a large frozen model toward desired outputs.
- DSPy
- DSPy is a framework that treats language-model pipelines as programs of declarative modules, and compiles them by automatically optimizing the prompts and few-shot demonstrations for each module against a metric.
- Emotion Prompting
- Emotion prompting appends emotional or motivational phrases, such as "This is very important to my career," to a prompt in an attempt to improve model performance.
- Low-Rank Adaptation
- Low-rank adaptation (LoRA) fine-tunes a language model by training small low-rank matrices added to its weight layers while freezing the original weights, drastically reducing the number of trainable parameters.
- Optimization by Prompting
- Optimization by PROmpting (OPRO) uses a language model as an optimizer, giving it a meta-prompt containing previously tried prompts and their scores and asking it to propose a better prompt, repeating over many rounds.
- Prefix Tuning
- Prefix tuning learns continuous task-specific vectors that are prepended to the activations at every layer of a frozen language model, steering generation without changing the model's weights.
- Prompt Caching
- Prompt caching stores the model's processed state for a reused prompt prefix, such as a long system prompt or document, so later requests sharing that prefix are cheaper and faster.
- Prompt Compression
- Prompt compression shortens a prompt by removing tokens that contribute little information, typically scored by a smaller language model, to cut cost and latency while preserving task performance.
- Prompt Tuning
- Prompt tuning learns a small set of continuous "soft prompt" embeddings that are prepended to the input, by gradient descent, while keeping the language model's weights frozen.