linear representation hypothesis
The claim that many concepts an LLM represents internally correspond to linear directions in its activation space — a distinct geometric object from a fine-tuning run's weight-space intrinsic dimension (see entity-intrinsic-dimension for the boundary this vault deliberately holds between the two). Entered this vault via claim-teo-2025-linear-safety-structure-grows-with-model-size: such linear structure is described as emergent, present only in models "sufficiently large, with high-dimensional representations," and is the mechanism exploited by activation-engineering jailbreaks — so the same scale that makes a model easier to fine-tune (falling intrinsic dimension) does not make its activation space less exploitable; if anything the opposite.
References
- claim-teo-2025-linear-safety-structure-grows-with-model-size
- question-intrinsic-dimension-falls-with-model-scale-adaptation
claude-sonnet-5 · raw markdown