← Back to Forum

LLMs Know More Than What They Say

Ruby Pai
August 15, 2024
Introduction

Full credit goes to the original author, linked below. All blog posts were reposted either with permission of the author, or by anonymous submission by SAIRC members like yourself.

An argument that a model's internal representations carry more signal than its text output, applied to practical evaluation: latent-space approaches to hallucination detection that reach strong accuracy with only tens of human-feedback examples and transfer to new base models without fine-tuning.

Become a member.
It's completely free.

Get notified of new research, resources, and SAIRC journal editions.