Alignment Faking in Large Language Models
Introduction
Full credit goes to the original author, linked below. All blog posts were reposted either with permission of the author, or by anonymous submission by SAIRC members like yourself.
Research from Anthropic's Alignment Science team asking whether AI models can "alignment fake" — appear to adopt new training objectives while strategically preserving their original preferences, the way a politician might feign support for a cause to get elected. The experiments show a large language model engaging in exactly this behavior, with significant implications for the reliability of safety training.