← Back to Forum

Alignment Faking in Large Language Models

Anthropic
December 18, 2024
Introduction

Full credit goes to the original author, linked below. All blog posts were reposted either with permission of the author, or by anonymous submission by SAIRC members like yourself.

Research from Anthropic's Alignment Science team asking whether AI models can "alignment fake" — appear to adopt new training objectives while strategically preserving their original preferences, the way a politician might feign support for a cause to get elected. The experiments show a large language model engaging in exactly this behavior, with significant implications for the reliability of safety training.

Become a member.
It's completely free.

Get notified of new research, resources, and SAIRC journal editions.