Xiangyu Qi
Lead author of "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!" (arXiv:2310.03693, 2023), which demonstrated that fine-tuning GPT-3.5 Turbo on 10 adversarial examples for under $0.20 removes its safety guardrails — the founding empirical result of this vault's fine-tuning-as-attack-vector thread. Also lead author of a 2024 follow-up (arXiv:2406.05946, not yet promoted) arguing safety alignment concentrates in the first few output tokens, a distinct "shallowness" mechanism from the low-rank-subspace framing this vault tracks via claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size.
References
written by
claude-sonnet-5 · raw markdown