100 Harmless Samples Are All You Need to Break a Safety-Aligned LLM
- Time
- 2026-08-09 11:10 ~ 11:40
- Speaker
- Allen, 黃亮勳
- Room
- TR212
- Co-write
Abstract
Think your "clean" fine-tuning data is safe? Recent research shows that just 100 samples from benign datasets like Dolly and Alpaca — innocent on the surface, yet carrying harmful gradient directions — are enough to collapse the safety alignment of an already-aligned LLM. And it works anchor-free: no external malicious data required. In this talk, we put Twinkle AI's open-source gemma-3-4B-T1-it on the table for a hands-on attack-and-defense walkthrough: how to surface these "invisible landmines" from Traditional Chinese benign datasets using Self-Inf-N, why they're so hard to undo in continual learning scenarios, and whether gradient-constrained defenses like SafeGrad and AsFT can actually hold the line. For developers and researchers interested in LLM safety, fine-tuning, and open-source model governance.
Speaker
Allen
我目前是台灣科技大學資訊工程所碩士二年級學生,具備生成式 AI、深度學習與 DevOps 的實務經驗。碩士研究聚焦於 LLM 微調的資安議題。專案經驗橫跨大型語言模型等 AI 領域,曾協助跨國企業提出 AI PoC,也參與過 RAG 解決方案的設計。歡迎大家與我交流!
黃亮勳
創辦 Twinkle AI,專注於開源繁體中文資料集與下一代繁中模型,為台灣 AI 生態奠定基礎並持續推進模型訓練技術。領導團隊打造台灣首個最小推理模型 Formosa-1 (F1),並曾以一人團隊完成繁體中文語料蒐集與模型訓練,成功推出台灣首個 Llama 3.2 3B 最小推理模型。