Feature learning, alignment and the linear representation hypothesis for steering and monitoring LLMs
A trained Large Language Model (LLM) contains much of human knowledge. Yet, it is difficult to gauge the extent or accuracy of that knowledge, as LLMs do not always ``know what they know'' and may even be unintentionally or actively misleading. In this talk I will discuss feature learning and some interesting behaviors of MLPs, seemingly related to LLMs. I will introduce Recursive Feature Machines—a powerful method originally designed for extracting relevant features from tabular data. I will show how this technique enables us to detect and guide LLM behaviors toward almost any desired concept by adding a multiple of fixed vectors in LLM activation spaces. Finally I will discuss a few, perhaps only tangentially related, thoughts on alignment.

