About Me

🎯I’m looking for full-time positions!

Hello, I am Tinghao Xie 谢廷浩, a final year ECE PhD candidate at Princeton, advised by Prof. Prateek Mittal. I was also a student researcher at TikTok and at Meta. Previously, I earned my Bachelor degree in Computer Science at Zhejiang University.

I work on AI safety and security. I am particularly interested in how models fail, how to evaluate them reliably, and how to improve their capabilities and safeguards -- across multimodal models and systems (current focus), language models, and learning systems.

Selected Research

📖 Red-teaming NSFW Image Classifiers as Text-to-Image Safeguards
Tinghao Xie, Yueqi Xie, Alireza Zareian, Shuming Hu, Felix Juefei-Xu, Xiaowen Lin, Ankit Jain, Prateek Mittal, Li Chen
ACL 2026 Findings

📖 SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors
Tinghao Xie*, Xiangyu Qi*, Yi Zeng*, Yangsibo Huang*, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, Prateek Mittal
ICLR 2025

📖 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Xiangyu Qi*, Yi Zeng*, Tinghao Xie*, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal$^†$, Peter Henderson$^†$
ICLR 2024 (oral)
📰 This work was exclusively reported by New York Times, and covered by many other social medias!

All publications

What I study

Selected directions & work
Current focus

Multimodal
models & systems

What models see.
What models know.
When we can trust them.

Visual knowledge: people, brands, films, landmarks, and animals.

Knowledge

Learning beyond what models know

Learn visual knowledge from people and real-world evidence, using VLMs to interpret and connect it.

Visual knowledge learningTikTok · Unpublished
Three image contexts keep the same red-framed target while the surrounding benign objects change.

Perception & safety

Seeing what safeguard VLMs miss

Reveal how visual context misleads safeguard VLMs, then improve their robustness.

Red-teaming visual safeguards Meta · ACL Findings ’26
Also exploringCopyright safeguards ICLR ’25 ↗Watermark removal attacks

How can we safely delegate more AI evaluation and training work to AI?

01People provide gold labels
A human annotator labels model responses Safe or Unsafe, according to whether they fulfill an unsafe request.

Human labels to validate LLM evaluators.

SORRY-Bench

Which safety evaluator should we trust?

Humans label whether model responses fulfill unsafe requests, giving us gold labels to compare model judges and their settings.

Reliable, but expensive to scale. This reliance on manual annotation motivates more scalable oversight methods.

ICLR ’25
02People guide more reliable AI data synthesis
A person designs prompt guidance to steer synthesis toward images with more reliable intended labels.

Steered synthesis to test T2I safety monitors.

Visual safeguards · Meta

How can we evaluate VLMs that judge T2I safety?

There may be no stronger judge we can trust to evaluate them. Simply prompting a diffusion model for unsafe images does not reliably produce unsafe images.

Steering data synthesis to be more reliable. Human-designed prompt suffixes made synthesizing unsafe images much more reliable, reducing the need for image-by-image annotation.

ACL Findings ’26
03AI curates visual knowledge from people
People create short videos and community posts, providing external evidence that a VLM interprets to curate visual knowledge data.

Human evidence for curating new visual knowledge.

Visual knowledge · TikTok

What if even the strongest model lacks the relevant visual knowledge?

We cannot distill knowledge a model lacks. Relevant evidence may instead be scattered across what people share on social media every day.

VLMs curate knowledge from human evidence. Use VLMs to interpret and connect people’s real-world experiences, curating visual knowledge data for further learning.

TikTok · Unpublished

So, what can we delegate to models, and how?

Reliable delegation still depends on people. Across these projects, human judgments help validate model evaluators, human-designed synthesis makes evaluation data more reliable, and people’s real-world experiences provide evidence for knowledge that even strong models lack.

As models take on more evaluation, training-data generation, and oversight (e.g., CoT monitoring), I want to understand when we can rely on them, and what human input or external evidence is still needed to verify and improve their outputs.