Interpretability

LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit

Anonymous
May 2, 2026
Abstract
This paper was either an anonymous submission of interesting research or was written by a student; full credit remains with the author(s), linked below. When a language model sycophantically agrees with a user's false belief, is it failing to detect the error, or noticing and agreeing anyway? We show the second. Across twelve open-weight models from five labs (1.5B–72B), the same small set of attention heads carries a “this statement is wrong” signal whether the model is evaluating an isolated claim or being pressured to agree with a user. Silencing these heads in Gemma-2-2B flips sycophancy from 28% to 81% while factual accuracy moves only from 69% to 70%; the circuit controls deference, not knowledge. Edge-level path patching confirms the same connections between heads span sycophancy, factual lying, and instructed lying (r>0.97 on Gemma-2-2B, r=0.988–0.995 on Phi-4). Opinion-agreement, where there is no factual ground truth, reuses these head positions but writes into an orthogonal direction, so the substrate is not a relabeled “truth direction.” Alignment training masks but does not remove this circuit: Meta's Llama-3.1→ 3.3 RLHF refresh cut sycophancy tenfold while the shared heads persisted and the projection-ablation effect grew (substrate persistence replicates on Mistral→ Zephyr at 7B, independent family), and our own anti-sycophancy DPO reduced sycophancy 46–93% on two models without moving probe transfer. When these models sycophant, they register the error and agree anyway.
Full Paper

Become a member.
It's completely free.

Get notified of new research, resources, and SAIRC journal editions.