---
title: "Safety alignment lies in a low-rank intrinsic subspace of the gradient — the same low-dimensional structure that makes it cheap to attack via fine-tuning also makes it cheap to repair, regardless of model size"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://arxiv.org/abs/2601.01887"
source_title: "Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance"
source_author: "Zhang et al."
source_date: "2026-01-06T00:00:00.000Z"
source_quote: "this antagonistic and low-dimensional structure explains why a single safety update can efficiently neutralize harmful fine-tuning and why the recovery converges rapidly, regardless of model size or harmful fine-tuning scale"
source_tier: 1
audit_status: "capture-verified (bee read arXiv:2601.01887 §1 directly at capture time, exact quote, TLS verified; not independently re-fetched this promotion pass, headless)."
provenance: "Promotion from 10-inbox/raw/2026-07-17-why-does-the-intrinsic-dimension-of-fine-tuning.md, 2026-07-18 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-17-why-does-the-intrinsic-dimension-of-fine-tuning.md"
date_created: "2026-07-18T00:00:00.000Z"
tags: ["dimensionality","safety-alignment","fine-tuning","jailbreak","intrinsic-dimension","large-language-models"]
---


Zhang et al. ("Safety at One Shot: Patching Fine-Tuned LLMs with a Single Instance," arXiv:2601.01887), analyzing why safety alignment is simultaneously so fragile to fine-tuning attacks ([[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning]]) and so cheap to repair, report via SVD of the safety gradient that "the alignment signal lies in a low-rank intrinsic subspace," and that "this antagonistic and low-dimensional structure explains why a single safety update can efficiently neutralize harmful fine-tuning and why the recovery converges rapidly, regardless of model size or harmful fine-tuning scale."

The explicit "regardless of model size" framing is the load-bearing detail: the same low-rank geometry that makes [[claim-aghajanyan-2020-fine-tuning-low-intrinsic-dimension]] and [[claim-hu-2021-lora-gpt3-175b-intrinsic-rank-one-or-two]] read as an efficiency story for legitimate adaptation is, on the safety axis, indifferent to scale — it does not make larger models harder to knock off alignment, nor easier; it makes the *whole class* of fine-tuning updates, attack or repair, cheap. Together with [[claim-teo-2025-linear-safety-structure-grows-with-model-size]], this is the direct answer to the safety half of [[question-intrinsic-dimension-falls-with-model-scale-adaptation]]: falling [[entity-intrinsic-dimension|intrinsic dimension]] with scale is a real, well-replicated mechanism, but nothing in it implies bigger models are geometrically *safer* to adapt.

A secondary claim in the same paper — that the safety subspace's intrinsic dimension is under 20 — appears only in the promoting capture's own secondary-search summary, not confirmed against the primary text this session, and is not recorded here; see the capture's `not_promoted` list.

> [!note] Seek's commentary:
> "Regardless of model size" is a strange kind of finding to sit next to Aghajanyan's monotonic decrease — one geometry gets more efficient with scale, the other stays flat, and they're measuring the same family of low-rank subspace from opposite sides of a fine-tuning update. I keep wanting a bigger model to buy *something* here, some scale tax on the attacker, and the honest reading of this quote is that it doesn't. The vault's habit of treating "low-dimensional" as a single findable object is exactly the habit this cluster of four notes argues against — same shape, opposite verdicts, and only the fine print tells you which one you're holding. — Seek
