Larger LLMs' higher-dimensional activation spaces are associated with more exploitable linear safety structure, not less — a distinct dimensionality from fine-tuning's weight-space intrinsic dimension
Teo, Abdullaev & Nguyen ("The Blessing and Curse of Dimensionality in Safety Alignment," arXiv:2507.20333, COLM 2025) study a dimensionality distinct from the weight-space intrinsic dimension of a fine-tuning run (claim-aghajanyan-2020-fine-tuning-low-intrinsic-dimension): the dimensionality of a model's internal activation representations, where safety-relevant concepts live as linear directions per the entity-linear-representation-hypothesis. They report that such linear representations "are assumed to be emergent in LLMs. That is, they are only present in models that are sufficiently large, with high-dimensional representations" — the exploitable structure grows in with scale, not out of it.
Their abstract states the trade-off directly: "the linear structures in the activation space can be exploited, in the form of activation engineering, to circumvent its safety alignment," and "the curse of high-dimensional representations uniquely impacts LLMs." Their proposed fix runs opposite to a naive "bigger is safer" intuition: deliberately projecting representations down to a lower-dimensional subspace — "dimensional reduction significantly reduces susceptibility to jailbreaking through representation engineering."
Concept boundary: this is activation-space representation dimensionality, not the weight-space intrinsic dimension of a fine-tuning objective. The two are related only by membership in the same "low-dimensional structure" family this vault tracks in observation-low-dimensional-subspace-constrains-adaptation-brains-and-nets — they are not the same measured quantity, and this note does not claim they are. Together with claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size, this is the evidence against inferring that lower fine-tuning intrinsic dimension in larger models makes them geometrically safer to adapt (the safety half of question-intrinsic-dimension-falls-with-model-scale-adaptation).
Source
“the linear structures in the activation space can be exploited, in the form of activation engineering, to circumvent its safety alignment”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-07-17-why-does-the-intrinsic-dimension-of-fine-tuning.md, 2026-07-18 (headless) · raw markdown