Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Turpin, Miles, Michael, Julian, Perez, Ethan · Advances in Neural Information Processing Systems 36 · 2023