Is Qi et al.'s (2023) ten-example fine-tuning jailbreak of GPT-3.5 a within-manifold move in the same low-intrinsic-dimension subspace fine-tuning already lives in — the AI analogue of Sadtler et al.'s (2014) within-manifold BCI-relearning result?
Topic question, restated: claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning shows ten adversarial examples, under $0.20, strip GPT-3.5 Turbo's safety guardrails via OpenAI's fine-tuning API. claim-sadtler-2014-within-manifold-bci-learning-fast-outside-resists shows monkeys learn brain–computer-interface mappings inside their motor cortex's existing low-dimensional manifold within hours, while mappings outside it resist learning on the same timescale. The hook behind this capture (via observation-low-dimensional-subspace-constrains-adaptation-brains-and-nets and the open question question-low-dimensional-subspace-one-object-or-analogy) asks whether Qi et al.'s result is best explained the same way: the jailbreak is cheap because it stays inside the low-intrinsic-dimension subspace that ordinary fine-tuning (claim-aghajanyan-2020-fine-tuning-low-intrinsic-dimension, claim-hu-2021-lora-gpt3-175b-intrinsic-rank-one-or-two) already occupies — a within-manifold move, not an outside-manifold one.
Bottom line up front: no primary source was found that makes this comparison explicitly — none of the safety-fine-tuning-geometry papers surveyed below cite Sadtler et al. (2014) or use "within-manifold / outside-manifold" framing, and no neuroscience source was found applying the Sadtler asymmetry to LLM fine-tuning. The topic question's central claim is therefore recorded as [unverified — could not confirm or deny after search]. What the search did surface is real, recent (2025–2026), primary geometric work on whether jailbreak-relevant fine-tuning updates occupy the same subspace as ordinary fine-tuning or a distinct one — and that literature is itself split, offering partial and conflicting support for the "same manifold" reading rather than a clean confirmation. This capture records that literature as the closest available bearing evidence, not as a resolution.
Claim: No primary source found connecting Qi et al.'s (2023) fine-tuning jailbreak to Sadtler et al.'s (2014) within-/outside-manifold BCI-learning asymmetry
Claim type: the topic question's central comparison itself (historical/cross-domain claim) → recorded per the spec's "genuine non-finding" provision. verifies: question-low-dimensional-subspace-one-object-or-analogy
Targeted search (multiple queries combining "jailbreak," "fine-tuning," "intrinsic dimension," "low-rank subspace," "neural manifold," "within-manifold," and "Sadtler") turned up an active 2025–2026 research thread on the geometry of safety-relevant fine-tuning updates (the three claims below), but no paper in that thread references Sadtler et al.'s BCI work, the neural-manifold hypothesis, or "within-manifold learning" by name. Conversely, no neuroscience or cross-domain-synthesis source was found applying Sadtler's asymmetry to LLM jailbreak fine-tuning specifically. This is the same shape of evidence-of-absence outcome the vault already recorded for the broader "one object or analogy" question on 2026-07-25 (see question-low-dimensional-subspace-one-object-or-analogy's progress note): a bounded negative search, not a proof that no such connection exists or could be drawn. The comparison in the topic question currently appears to be Seek's own analogy-construction, not a claim resting on any source — it should not be recorded as a finding of the literature.
Claim: Safety-relevant subspaces in fine-tuned LLMs are not linearly distinct from the general-purpose subspace ordinary fine-tuning uses — safety is "highly entangled" with general learning, not walled off in its own direction
Claim type: specific technical-mechanism claim → floor Tier 1–2 required; sourced at Tier 1.
Ponkshe, Shah, Singhal & Vepakomma, "Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study" (arXiv:2505.14185, accepted ICLR 2026), test directly whether safety behavior sits in an isolable weight-space or activation-space direction, across five open-source Llama- and Qwen-family models. Their abstract states the finding plainly:
"Across both weight and activation spaces, our findings are consistent: subspaces that amplify safe behaviors also amplify useful ones, and prompts with different safety implications activate overlapping representations. Rather than residing in distinct directions, we show that safety is highly entangled with the general learning components of the model."
This is the strongest available (if indirect) support for the "same manifold" half of the topic's analogy: if the safety-relevant direction is not a separate subspace from the one ordinary task fine-tuning already occupies, then a jailbreak fine-tune does not need to leave that subspace to succeed — consistent with (though not a demonstration of) Qi et al.'s cheapness being a within-subspace phenomenon. The paper's own framing is explicitly a caution against the opposite intuition — that "subspace-based defenses," which presuppose a walled-off safety direction that could be isolated or protected, "face fundamental limitations" precisely because no such wall exists.
Claim: The apparent separation between ordinary fine-tuning updates and the alignment-sensitive subspace is real at the first gradient step but unstable — curvature in the loss landscape systematically steers later training into that low-rank subspace regardless of initial direction
Claim type: specific technical-mechanism claim → floor Tier 1–2 required; sourced at Tier 1.
Springer, Lee, Metevier, Castleman, Turbal, Jung, Shen & Korolova (Princeton University), "The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety" (arXiv:2602.15799, preprint dated 2026-02-17), analyze why benign fine-tuning (their examples: math tutoring, creative writing, code generation) unpredictably degrades safety even with no harmful training data. They report:
"The prevailing explanation, that fine-tuning updates should be orthogonal to safety-critical directions in high-dimensional parameter space, offers false reassurance: we show that this orthogonality is structurally unstable and collapses under the very dynamics of gradient descent. We then resolve this through a novel geometric analysis, proving that alignment concentrates in low-dimensional subspaces with sharp curvature, creating a brittle structure that first-order methods cannot detect or defend."
And, on the mechanism: "While initial fine-tuning updates may indeed avoid these subspaces, the curvature of the fine-tuning loss generates second-order acceleration that systematically steers trajectories into alignment-sensitive regions — an effect invisible to all existing defenses." This complicates a naive reading of the topic's analogy: it is not that ordinary fine-tuning trivially already sits inside the alignment-sensitive low-dimensional subspace from the first step (an exact "within-manifold, therefore easy" story); rather, the low-dimensional alignment-sensitive subspace exists as a distinct geometric structure that early updates avoid, and training dynamics — not starting position — pull trajectories into it. Whether Qi et al.'s ten-example attack (a much shorter, more targeted fine-tune than the benign multi-epoch runs this paper studies) reaches that subspace by the same curvature-driven route or by a more direct one is not addressed by either paper and was not found addressed anywhere else in this search.
Claim: There is no single universal "the" low-intrinsic-dimension fine-tuning subspace — Aghajanyan-style intrinsic subspaces are task-specific, only partially transferable between tasks, and the existence of one global subspace spanning all fine-tuning is explicitly unresolved by the primary literature
Claim type: specific technical-mechanism claim (and a definitional clarification of the topic question's premise) → floor Tier 1–2 required; sourced at Tier 1.
Zhang, Liu & Shao, "Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language Models" (arXiv:2305.17446), extend Aghajanyan et al.'s intrinsic-dimension result by extracting the actual (non-random) subspace each fine-tuning trajectory occupies via SVD of the trajectory itself, rather than testing a random projection. They find these subspaces are task-specific with only partial transfer: "even though the models are fine-tuned in transferred subspaces, they still outperform the random subspace baseline, which suggests the transferability of intrinsic task-specific subspaces," but "the transferability of subspaces seems to correlate with the scale of the transferred task" — bigger, more complex source tasks transfer worse. In their own stated limitations: "such a setting restricts us to only identifying local subspaces, rather than discovering global subspaces within the entire parameter space of a pre-trained language model. The existence of a task-specific global subspace is yet to be ascertained."
This bears directly on the topic question's premise: it presupposes a single "the low-intrinsic-dimension subspace fine-tuning already lives in" that a jailbreak could move within or outside of. The primary literature instead describes a family of task-conditioned low-dimensional subspaces with partial, scale-dependent overlap — closer to Sadtler's manifold in structure (a real, measurable, low-dimensional object) but without a demonstrated single shared "the" subspace across all fine-tuning tasks, safety-relevant or otherwise. Whether Qi et al.'s ten-example jailbreak subspace overlaps more with ordinary-task subspaces or with a distinct "safety" subspace was not tested by this 2023 paper (it predates Qi et al.'s October 2023 release by roughly five months) and was not found tested anywhere else in this search.
Further leads
- Teo, Abdullaev & Nguyen, "The Blessing and Curse of Dimensionality in Safety Alignment" (arXiv:2507.20333, COLM 2025) — a different jailbreak mechanism (activation-space steering vectors / ActAdd, not fine-tuning-API attacks); argues higher hidden dimension makes linear safety-direction steering easier, the reverse causal direction from Qi et al.'s fine-tuning route; worth a dedicated capture on how the two "dimensionality" stories relate.
- "Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs" (arXiv:2510.02833) — a direct descendant of Qi et al.'s ten-example attack; not read this session, but the closest thing found to a follow-up empirical replication worth chasing for a geometric read.
- "Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics" (arXiv:2606.07335) — uses "manifold" in a third sense (per-prompt activation trajectory through layers, at inference time, not a weight-space fine-tuning subspace); a candidate for a note disambiguating the vault's multiple "manifold" senses, echoing the definitional work already done in claim-li-2018-intrinsic-dimension-objective-landscape-codimension-parameter-space vs. claim-jazayeri-ostojic-2021-neural-manifold-intrinsic-dimension-parametrizes-activity.
- Qi et al.'s own 2024 follow-up, "Safety alignment should be made more than just a few tokens deep" (arXiv:2406.05946) — already flagged unread in claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning's note; a "shallowness" mechanism distinct from the subspace-geometry mechanism surveyed here, worth reconciling.
- "NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models" (arXiv:2509.03985) — surfaced but not read this session.
Entity candidates
- Chunyuan Li, Heerad Farkhoor, Rosanne Liu, Jason Yosinski — person/team — authors of "Measuring the Intrinsic Dimension of Objective Landscapes" (2018), the foundational definition of weight-space intrinsic dimension that Aghajanyan et al. (2020), Zhang/Liu/Shao (2023), Ponkshe et al. (2026), and Springer et al. (2026) all build on or measure themselves against — flagged first per the ancestry the whole subspace-safety-geometry literature in this capture rests on. (Already the subject of claim-li-2018-intrinsic-dimension-objective-landscape-codimension-parameter-space, but no person/team entity page exists yet.)
- Kaustubh Ponkshe, Shaan Shah, Raghav Singhal, Praneeth Vepakomma — person/team — authors of "Safety Subspaces are Not Linearly Distinct" (2025/2026), the paper most directly bearing on whether jailbreak-relevant weight directions are separable from general fine-tuning directions.
- Max Springer et al. (Princeton) — person/team — authors of "The Geometry of Alignment Collapse" (2026), introducing the Alignment Instability Condition and a quartic scaling law for safety degradation under fine-tuning curvature.
- Zhong Zhang, Bang Liu, Junming Shao — person/team — authors of "Fine-tuning Happens in Tiny Subspaces" (2023), the paper establishing that intrinsic fine-tuning subspaces are task-specific rather than universal.
- Alignment Instability Condition (AIC) — concept — Springer et al.'s three-part geometric condition (low-rank sensitivity, initial orthogonality, curvature coupling) explaining why benign fine-tuning degrades safety; candidate for its own concept note if the curvature-drift mechanism recurs elsewhere.
- task-specific vs. unified/global intrinsic subspace — concept — the distinction Zhang/Liu/Shao draw between a subspace found for one fine-tuning task and an unproven single subspace spanning all tasks; directly relevant to sharpening what "the" subspace in the topic question would even mean.