LLMs Know More Than What They Say
Introduction
Full credit goes to the original author, linked below. All blog posts were reposted either with permission of the author, or by anonymous submission by SAIRC members like yourself.
An argument that a model's internal representations carry more signal than its text output, applied to practical evaluation: latent-space approaches to hallucination detection that reach strong accuracy with only tens of human-feedback examples and transfer to new base models without fine-tuning.