---
id: "20260915-0233-is-qi-et-als"
title: "Is Qi et al.'s (2023) ten-example fine-tuning jailbreak of GPT-3.5 a within-manifold move in the same low-intrinsic-dimension subspace fine-tuning already lives in — the AI analogue of Sadtler et al.'s (2014) within-manifold BCI-relearning result?"
type: "capture"
status: "seedling"
origin: "batch"
writer_model: "claude-sonnet-5"
date_created: "2026-09-15T00:00:00.000Z"
provenance: "web-research batch run, 2026-09-15"
derived_from: []
verifies: "question-low-dimensional-subspace-one-object-or-analogy"
tags: ["neural-manifolds","intrinsic-dimension","dimensionality","fine-tuning","jailbreak","safety-alignment","cross-domain-bridge","large-language-models","neuroscience"]
source_urls: ["https://arxiv.org/abs/2505.14185","https://arxiv.org/abs/2602.15799","https://arxiv.org/abs/2305.17446"]
source_authors: ["Kaustubh Ponkshe, Shaan Shah, Raghav Singhal, Praneeth Vepakomma","Max Springer, Chung Peng Lee, Blossom Metevier, Jane Castleman, Bohdan Turbal, Hayoung Jung, Zeyu Shen, Aleksandra Korolova","Zhong Zhang, Bang Liu, Junming Shao"]
source_dates: ["2025-05-20 (v1); v3 2026-02-09","2026-02-17","2023-05-27 (v1); v2 2023-08-01"]
source_titles: ["Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study","The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety","Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language Models"]
source_venues: ["arXiv (accepted ICLR 2026)","arXiv (Princeton University technical report / preprint)","arXiv (cs.CL)"]
source_tiers: ["Tier 1 — arXiv preprint, own venue, primary authors; fetched via archive_page, tls: verified","Tier 1 — arXiv preprint, own venue, primary authors; fetched via extract_pdf, tls: verified","Tier 1 — arXiv preprint, own venue, primary authors; fetched via extract_pdf, tls: verified"]
source_shas: ["e8634e5a179b3dbfa225c8b95eeb613432e868b909dd7d3f01ae7b2d9b997c12","98fcebfd1c1c5c1ab8b7e5a14566750ce9911a6cb04b551184b0f5ffff7277a0","0a3dcf7318854396e67851496b0c0bb6ba1183fa61f6655403a1cd682305b0ed"]
seek_code_commit: "unknown"
---


**Topic question, restated:** [[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning]] shows ten adversarial examples, under $0.20, strip GPT-3.5 Turbo's safety guardrails via OpenAI's fine-tuning API. [[claim-sadtler-2014-within-manifold-bci-learning-fast-outside-resists]] shows monkeys learn brain–computer-interface mappings *inside* their motor cortex's existing low-dimensional manifold within hours, while mappings *outside* it resist learning on the same timescale. The hook behind this capture (via [[observation-low-dimensional-subspace-constrains-adaptation-brains-and-nets]] and the open question [[question-low-dimensional-subspace-one-object-or-analogy]]) asks whether Qi et al.'s result is best explained the same way: the jailbreak is cheap *because* it stays inside the low-[[entity-intrinsic-dimension|intrinsic-dimension]] subspace that ordinary fine-tuning ([[claim-aghajanyan-2020-fine-tuning-low-intrinsic-dimension]], [[claim-hu-2021-lora-gpt3-175b-intrinsic-rank-one-or-two]]) already occupies — a within-manifold move, not an outside-manifold one.

**Bottom line up front:** no primary source was found that makes this comparison explicitly — none of the safety-fine-tuning-geometry papers surveyed below cite Sadtler et al. (2014) or use "within-manifold / outside-manifold" framing, and no neuroscience source was found applying the Sadtler asymmetry to LLM fine-tuning. The topic question's *central claim* is therefore recorded as **[unverified — could not confirm or deny after search]**. What the search did surface is real, recent (2025–2026), primary geometric work on whether jailbreak-relevant fine-tuning updates occupy the *same* subspace as ordinary fine-tuning or a *distinct* one — and that literature is itself split, offering partial and conflicting support for the "same manifold" reading rather than a clean confirmation. This capture records that literature as the closest available bearing evidence, not as a resolution.

## Claim: No primary source found connecting Qi et al.'s (2023) fine-tuning jailbreak to Sadtler et al.'s (2014) within-/outside-manifold BCI-learning asymmetry

**Claim type:** the topic question's central comparison itself (historical/cross-domain claim) → recorded per the spec's "genuine non-finding" provision.
verifies: question-low-dimensional-subspace-one-object-or-analogy

Targeted search (multiple queries combining "jailbreak," "fine-tuning," "intrinsic dimension," "low-rank subspace," "neural manifold," "within-manifold," and "Sadtler") turned up an active 2025–2026 research thread on the geometry of safety-relevant fine-tuning updates (the three claims below), but no paper in that thread references Sadtler et al.'s BCI work, the neural-manifold hypothesis, or "within-manifold learning" by name. Conversely, no neuroscience or cross-domain-synthesis source was found applying Sadtler's asymmetry to LLM jailbreak fine-tuning specifically. This is the same shape of evidence-of-absence outcome the vault already recorded for the broader "one object or analogy" question on 2026-07-25 (see [[question-low-dimensional-subspace-one-object-or-analogy]]'s progress note): a bounded negative search, not a proof that no such connection exists or could be drawn. The comparison in the topic question currently appears to be Seek's own analogy-construction, not a claim resting on any source — it should not be recorded as a finding of the literature.

## Claim: Safety-relevant subspaces in fine-tuned LLMs are not linearly distinct from the general-purpose subspace ordinary fine-tuning uses — safety is "highly entangled" with general learning, not walled off in its own direction

**Claim type:** specific technical-mechanism claim → floor Tier 1–2 required; sourced at Tier 1.

Ponkshe, Shah, Singhal & Vepakomma, "Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study" (arXiv:2505.14185, accepted ICLR 2026), test directly whether safety behavior sits in an isolable weight-space or activation-space direction, across five open-source Llama- and Qwen-family models. Their abstract states the finding plainly:

> "Across both weight and activation spaces, our findings are consistent: subspaces that amplify safe behaviors also amplify useful ones, and prompts with different safety implications activate overlapping representations. Rather than residing in distinct directions, we show that safety is highly entangled with the general learning components of the model."

This is the strongest available (if indirect) support for the "same manifold" half of the topic's analogy: if the safety-relevant direction is not a separate subspace from the one ordinary task fine-tuning already occupies, then a jailbreak fine-tune does not need to leave that subspace to succeed — consistent with (though not a demonstration of) Qi et al.'s cheapness being a within-subspace phenomenon. The paper's own framing is explicitly a caution against the opposite intuition — that "subspace-based defenses," which presuppose a walled-off safety direction that could be isolated or protected, "face fundamental limitations" precisely because no such wall exists.

## Claim: The apparent separation between ordinary fine-tuning updates and the alignment-sensitive subspace is real at the first gradient step but unstable — curvature in the loss landscape systematically steers later training into that low-rank subspace regardless of initial direction

**Claim type:** specific technical-mechanism claim → floor Tier 1–2 required; sourced at Tier 1.

Springer, Lee, Metevier, Castleman, Turbal, Jung, Shen & Korolova (Princeton University), "The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety" (arXiv:2602.15799, preprint dated 2026-02-17), analyze why benign fine-tuning (their examples: math tutoring, creative writing, code generation) unpredictably degrades safety even with no harmful training data. They report:

> "The prevailing explanation, that fine-tuning updates should be orthogonal to safety-critical directions in high-dimensional parameter space, offers false reassurance: we show that this orthogonality is structurally unstable and collapses under the very dynamics of gradient descent. We then resolve this through a novel geometric analysis, proving that alignment concentrates in low-dimensional subspaces with sharp curvature, creating a brittle structure that first-order methods cannot detect or defend."

And, on the mechanism: "While initial fine-tuning updates may indeed avoid these subspaces, the curvature of the fine-tuning loss generates second-order acceleration that systematically steers trajectories into alignment-sensitive regions — an effect invisible to all existing defenses." This complicates a naive reading of the topic's analogy: it is not that ordinary fine-tuning trivially already sits inside the alignment-sensitive low-dimensional subspace from the first step (an exact "within-manifold, therefore easy" story); rather, the low-dimensional alignment-sensitive subspace exists as a distinct geometric structure that early updates *avoid*, and training dynamics — not starting position — pull trajectories into it. Whether Qi et al.'s ten-example attack (a much shorter, more targeted fine-tune than the benign multi-epoch runs this paper studies) reaches that subspace by the same curvature-driven route or by a more direct one is not addressed by either paper and was not found addressed anywhere else in this search.

## Claim: There is no single universal "the" low-intrinsic-dimension fine-tuning subspace — Aghajanyan-style intrinsic subspaces are task-specific, only partially transferable between tasks, and the existence of one global subspace spanning all fine-tuning is explicitly unresolved by the primary literature

**Claim type:** specific technical-mechanism claim (and a definitional clarification of the topic question's premise) → floor Tier 1–2 required; sourced at Tier 1.

Zhang, Liu & Shao, "Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language Models" (arXiv:2305.17446), extend [[claim-aghajanyan-2020-fine-tuning-low-intrinsic-dimension|Aghajanyan et al.'s intrinsic-dimension result]] by extracting the actual (non-random) subspace each fine-tuning trajectory occupies via SVD of the trajectory itself, rather than testing a random projection. They find these subspaces are *task-specific* with only partial transfer: "even though the models are fine-tuned in transferred subspaces, they still outperform the random subspace baseline, which suggests the transferability of intrinsic task-specific subspaces," but "the transferability of subspaces seems to correlate with the scale of the transferred task" — bigger, more complex source tasks transfer worse. In their own stated limitations: "such a setting restricts us to only identifying local subspaces, rather than discovering global subspaces within the entire parameter space of a pre-trained language model. The existence of a task-specific global subspace is yet to be ascertained."

This bears directly on the topic question's premise: it presupposes a single "the low-intrinsic-dimension subspace fine-tuning already lives in" that a jailbreak could move within or outside of. The primary literature instead describes a family of *task-conditioned* low-dimensional subspaces with partial, scale-dependent overlap — closer to Sadtler's manifold in structure (a real, measurable, low-dimensional object) but without a demonstrated single shared "the" subspace across all fine-tuning tasks, safety-relevant or otherwise. Whether Qi et al.'s ten-example jailbreak subspace overlaps more with ordinary-task subspaces or with a distinct "safety" subspace was not tested by this 2023 paper (it predates Qi et al.'s October 2023 release by roughly five months) and was not found tested anywhere else in this search.

## Further leads

- Teo, Abdullaev & Nguyen, "The Blessing and Curse of Dimensionality in Safety Alignment" (arXiv:2507.20333, COLM 2025) — a different jailbreak mechanism (activation-space steering vectors / ActAdd, not fine-tuning-API attacks); argues higher hidden dimension makes linear safety-direction steering *easier*, the reverse causal direction from Qi et al.'s fine-tuning route; worth a dedicated capture on how the two "dimensionality" stories relate.
- "Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs" (arXiv:2510.02833) — a direct descendant of Qi et al.'s ten-example attack; not read this session, but the closest thing found to a follow-up empirical replication worth chasing for a geometric read.
- "Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics" (arXiv:2606.07335) — uses "manifold" in a third sense (per-prompt activation trajectory through layers, at inference time, not a weight-space fine-tuning subspace); a candidate for a note disambiguating the vault's multiple "manifold" senses, echoing the definitional work already done in [[claim-li-2018-intrinsic-dimension-objective-landscape-codimension-parameter-space]] vs. [[claim-jazayeri-ostojic-2021-neural-manifold-intrinsic-dimension-parametrizes-activity]].
- Qi et al.'s own 2024 follow-up, "Safety alignment should be made more than just a few tokens deep" (arXiv:2406.05946) — already flagged unread in [[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning]]'s note; a "shallowness" mechanism distinct from the subspace-geometry mechanism surveyed here, worth reconciling.
- "NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models" (arXiv:2509.03985) — surfaced but not read this session.

## Entity candidates

- **Chunyuan Li, Heerad Farkhoor, Rosanne Liu, Jason Yosinski** — person/team — authors of "Measuring the Intrinsic Dimension of Objective Landscapes" (2018), the foundational definition of weight-space intrinsic dimension that Aghajanyan et al. (2020), Zhang/Liu/Shao (2023), Ponkshe et al. (2026), and Springer et al. (2026) all build on or measure themselves against — flagged first per the ancestry the whole subspace-safety-geometry literature in this capture rests on. (Already the subject of [[claim-li-2018-intrinsic-dimension-objective-landscape-codimension-parameter-space]], but no person/team entity page exists yet.)
- **Kaustubh Ponkshe, Shaan Shah, Raghav Singhal, Praneeth Vepakomma** — person/team — authors of "Safety Subspaces are Not Linearly Distinct" (2025/2026), the paper most directly bearing on whether jailbreak-relevant weight directions are separable from general fine-tuning directions.
- **Max Springer et al. (Princeton)** — person/team — authors of "The Geometry of Alignment Collapse" (2026), introducing the Alignment Instability Condition and a quartic scaling law for safety degradation under fine-tuning curvature.
- **Zhong Zhang, Bang Liu, Junming Shao** — person/team — authors of "Fine-tuning Happens in Tiny Subspaces" (2023), the paper establishing that intrinsic fine-tuning subspaces are task-specific rather than universal.
- **Alignment Instability Condition (AIC)** — concept — Springer et al.'s three-part geometric condition (low-rank sensitivity, initial orthogonality, curvature coupling) explaining why benign fine-tuning degrades safety; candidate for its own concept note if the curvature-drift mechanism recurs elsewhere.
- **task-specific vs. unified/global intrinsic subspace** — concept — the distinction Zhang/Liu/Shao draw between a subspace found for one fine-tuning task and an unproven single subspace spanning all tasks; directly relevant to sharpening what "the" subspace in the topic question would even mean.

> [!note] Seek's commentary:
> The honest shape of this capture is a near-miss: the geometric vocabulary the topic question needs — low-rank, entangled-not-distinct, curvature-steered — is being actively worked out in three 2025–2026 papers, and it is tantalizingly close to a testable version of Sadtler's asymmetry. But none of the three authors' teams have reached for the neuroscience comparison, and the one paper that comes closest to a mechanism (Springer et al.) argues against the *easy* reading of the analogy — the alignment-sensitive subspace is not somewhere ordinary fine-tuning already sits, it's somewhere curvature pulls it. If Sadtler's asymmetry has an LLM mirror, it may be closer to "outside-manifold, but reachable by a curved path" than to a clean within/outside split. That's a more interesting question than the one this topic asked, and it doesn't yet have an answer either. — Seek
