About Me
Hello, I am Tinghao Xie 谢廷浩, a final year ECE PhD candidate at Princeton, advised by Prof. Prateek Mittal. I was also a student researcher at TikTok and at Meta. Previously, I earned my Bachelor degree in Computer Science at Zhejiang University.
I work on AI safety and security. I am particularly interested in how models fail, how to evaluate them reliably, and how to improve their capabilities and safeguards -- across multimodal models and systems (current focus), language models, and learning systems.
Selected Research
📖 Red-teaming NSFW Image Classifiers as Text-to-Image Safeguards
Tinghao Xie, Yueqi Xie, Alireza Zareian, Shuming Hu, Felix Juefei-Xu, Xiaowen Lin, Ankit Jain, Prateek Mittal, Li Chen
ACL 2026 Findings
📖 SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors
Tinghao Xie*, Xiangyu Qi*, Yi Zeng*, Yangsibo Huang*, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, Prateek Mittal
ICLR 2025
📖 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Xiangyu Qi*, Yi Zeng*, Tinghao Xie*, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal$^†$, Peter Henderson$^†$
ICLR 2024 (oral)
📰 This work was exclusively reported by New York Times, and covered by many other social medias!
What I study
Multimodal
models & systems
What models see.
What models know.
When we can trust them.
Knowledge
Learning beyond what models know
Learn visual knowledge from people and real-world evidence, using VLMs to interpret and connect it.
Perception & safety
Seeing what safeguard VLMs miss
Reveal how visual context misleads safeguard VLMs, then improve their robustness.
Red-teaming visual safeguards Meta · ACL Findings ’26Language models
How adaptation changes safety, how safety works inside models, and how to measure it reliably.
Learning systems
Poisoning and backdoor threats across training and deployment.
How I work
A cycle of discovery and improvement.
Secure AI
Failures inform improvements.
Improvements create new questions to test.
01 / Research approach
Attack
Where do our assumptions break?
Stress-test models and safeguards to expose failures that ordinary use can miss.
How can we safely delegate more AI evaluation and training work to AI?
Human labels to validate LLM evaluators.
SORRY-Bench
Which safety evaluator should we trust?
Humans label whether model responses fulfill unsafe requests, giving us gold labels to compare model judges and their settings.
Reliable, but expensive to scale. This reliance on manual annotation motivates more scalable oversight methods.
ICLR ’25Steered synthesis to test T2I safety monitors.
Visual safeguards · Meta
How can we evaluate VLMs that judge T2I safety?
There may be no stronger judge we can trust to evaluate them. Simply prompting a diffusion model for unsafe images does not reliably produce unsafe images.
Steering data synthesis to be more reliable. Human-designed prompt suffixes made synthesizing unsafe images much more reliable, reducing the need for image-by-image annotation.
ACL Findings ’26Human evidence for curating new visual knowledge.
Visual knowledge · TikTok
What if even the strongest model lacks the relevant visual knowledge?
We cannot distill knowledge a model lacks. Relevant evidence may instead be scattered across what people share on social media every day.
VLMs curate knowledge from human evidence. Use VLMs to interpret and connect people’s real-world experiences, curating visual knowledge data for further learning.
TikTok · UnpublishedSo, what can we delegate to models, and how?
Reliable delegation still depends on people. Across these projects, human judgments help validate model evaluators, human-designed synthesis makes evaluation data more reliable, and people’s real-world experiences provide evidence for knowledge that even strong models lack.
As models take on more evaluation, training-data generation, and oversight (e.g., CoT monitoring), I want to understand when we can rely on them, and what human input or external evidence is still needed to verify and improve their outputs.