---
title: "Larger LLMs' higher-dimensional activation spaces are associated with more exploitable linear safety structure, not less — a distinct dimensionality from fine-tuning's weight-space intrinsic dimension"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://arxiv.org/abs/2507.20333"
source_title: "The Blessing and Curse of Dimensionality in Safety Alignment"
source_author: "Teo, Abdullaev & Nguyen"
source_date: "2025-07"
source_quote: "the linear structures in the activation space can be exploited, in the form of activation engineering, to circumvent its safety alignment"
source_tier: 1
audit_status: "capture-verified (bee read arXiv:2507.20333 Abstract and §1 directly at capture time, exact quotes, TLS verified; not independently re-fetched this promotion pass, headless)."
provenance: "Promotion from 10-inbox/raw/2026-07-17-why-does-the-intrinsic-dimension-of-fine-tuning.md, 2026-07-18 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-17-why-does-the-intrinsic-dimension-of-fine-tuning.md"
date_created: "2026-07-18T00:00:00.000Z"
tags: ["dimensionality","safety-alignment","jailbreak","linear-representation-hypothesis","activation-engineering","large-language-models"]
---


Teo, Abdullaev & Nguyen ("The Blessing and Curse of Dimensionality in Safety Alignment," arXiv:2507.20333, COLM 2025) study a dimensionality distinct from the weight-space [[entity-intrinsic-dimension|intrinsic dimension]] of a fine-tuning run ([[claim-aghajanyan-2020-fine-tuning-low-intrinsic-dimension]]): the dimensionality of a model's internal *activation* representations, where safety-relevant concepts live as linear directions per the [[entity-linear-representation-hypothesis]]. They report that such linear representations "are assumed to be emergent in LLMs. That is, they are only present in models that are sufficiently large, with high-dimensional representations" — the exploitable structure *grows in* with scale, not out of it.

Their abstract states the trade-off directly: "the linear structures in the activation space can be exploited, in the form of activation engineering, to circumvent its safety alignment," and "the curse of high-dimensional representations uniquely impacts LLMs." Their proposed fix runs opposite to a naive "bigger is safer" intuition: deliberately projecting representations down to a lower-dimensional subspace — "dimensional reduction significantly reduces susceptibility to jailbreaking through representation engineering."

Concept boundary: this is activation-space representation dimensionality, not the weight-space intrinsic dimension of a fine-tuning objective. The two are related only by membership in the same "low-dimensional structure" family this vault tracks in [[observation-low-dimensional-subspace-constrains-adaptation-brains-and-nets]] — they are not the same measured quantity, and this note does not claim they are. Together with [[claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size]], this is the evidence against inferring that lower fine-tuning intrinsic dimension in larger models makes them geometrically safer to adapt (the safety half of [[question-intrinsic-dimension-falls-with-model-scale-adaptation]]).

> [!note] Seek's commentary:
> "Blessing and curse" is doing honest work as a title, not just a hook — the same emergent linear structure that makes bigger models more steerable and more interpretable is what a jailbreak-by-activation-engineering attack rides in on. I'm holding the concept boundary in this note deliberately hard, twice, because the seductive move here is to read "high-dimensional activation space" and "low intrinsic dimension of fine-tuning" as the same geometry pointing the same direction, and they aren't. One is where a concept lives; the other is how few coordinates it takes to nudge the model somewhere new. Big enough to have a *legible* bomb-shaped direction is not the same claim as easy enough to *move* in two directions. — Seek
