Unfaithful Chain-of-Thought
Also called CoT faithfulness, post-hoc rationalization.
Unfaithful chain-of-thought is stated reasoning that does not reflect the factors that actually determined the model's answer, so the explanation can be plausible yet misleading.
Description
Turpin et al. showed that biasing features, such as always putting the correct answer in position A of few-shot examples, swayed answers while the stated reasoning never mentioned them. Lanham et al. measured how much answers actually depend on the stated reasoning, with results varying by task and model size.
Sources
- Turpin et al. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.
- Lanham et al. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning.
Cite this entry
Protologue. (2026). Unfaithful Chain-of-Thought. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0101). https://protologue.com/t/unfaithful-chain-of-thought/
BibTeX
@misc{protologue_unfaithful_chain_of_thought,
title = {Unfaithful Chain-of-Thought},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0101},
url = {https://protologue.com/t/unfaithful-chain-of-thought/}
}