About Me
Hello, I am Tinghao Xie 谢廷浩, a final year ECE PhD candidate at Princeton, advised by Prof. Prateek Mittal. I was also a student researcher at TikTok and at Meta. Previously, I earned my Bachelor degree in Computer Science at Zhejiang University.
- I analyze and attack safety / alignment / watermark mechanisms in AI systems built upon LLMs, VLMs, and T2I models.
- breaking security of text-to-image systems from end to end
- benchmarking safety refusal & safety durability of LLMs
- revealing how fine-tuning LLMs can compromise safety
- eliciting copyrighted content, removing AIGC watermarks, etc.
- I aim to build safer & more trustworthy foundation models + securer AI systems.
- mid-training VLMs to enhance visual knowledge
- fine-tuning VLMs for more robust safety detection
- securing AI systems against data poisoning and backdoor attacks
Selected Research
📖 Red-teaming NSFW Image Classifiers as Text-to-Image Safeguards
Tinghao Xie, Yueqi Xie, Alireza Zareian, Shuming Hu, Felix Juefei-Xu, Xiaowen Lin, Ankit Jain, Prateek Mittal, Li Chen
ACL 2026 Findings
📖 SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors
Tinghao Xie*, Xiangyu Qi*, Yi Zeng*, Yangsibo Huang*, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, Prateek Mittal
ICLR 2025
📖 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Xiangyu Qi*, Yi Zeng*, Tinghao Xie*, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal$^†$, Peter Henderson$^†$
ICLR 2024 (oral)
📰 This work was exclusively reported by New York Times, and covered by many other social medias!