Unsafe 2 Safe Controllable Image Anonymization for Downstream Utility

Dartmouth College
CVPR 2026

✔ Structurally Consistent Identity Privacy

Unsafe input image Privacy-preserving safe output image

✔ Demographic Neutralization

Unsafe input image Privacy-preserving safe output image

✔ Non-human Content Anonymization

Unsafe input image Privacy-preserving safe output image

Examples from Unsafe2Safe (U2S). For each case, the model converts an unsafe image into a privacy-preserving safe version. Examples demonstrate key capabilities that may appear simultaneously: (1) structure-preserving full body anonymization, (2) demographic neutralization (race entropy ↑), and (3) obfuscation of non-human confidential details.

Abstract

Large image datasets often contain privacy-sensitive content, and deep models trained on them can memorize and leak identifiable information. We introduce Unsafe2Safe, a scalable two-stage framework that converts raw image datasets into anonymized yet utility-preserving versions. In Stage 1, a vision-language model detects privacy risks and produces a privacy-safe public caption together with an LLM-generated edit instruction. In Stage 2, a diffusion editor uses these dual prompts to rewrite only identity-revealing regions while preserving structure and task-relevant semantics.

To evaluate anonymization effectiveness, we propose a unified suite of Quality, Cheating, Privacy, and Utility scores. Across Caltech101 and MIT Indoor67, Unsafe2Safe substantially reduces face similarity, text similarity, and demographic predictability while maintaining downstream accuracy comparable to models trained on raw data. Fine-tuning instruction-following editors on our generated pairs further improves both privacy and utility. Unsafe2Safe provides a practical path toward constructing large-scale, privacy-safe datasets without compromising visual semantics or model performance.

Method

Unsafe2Safe two-stage method overview
A VLM inspects the image for privacy risks. For flagged images, it generates a private caption and a public caption without sensitive details. An LLM then produces an edit instruction on how sensitive attributes should be modified. A diffusion editor uses these priors to generate a privacy-safe image while preserving scene semantics.
Qualitative comparison of Unsafe2Safe image edits
Qualitative comparison across model variants on MIT Indoor-67 scenes. All methods generally preserve the original content and spatial composition while applying varying levels of stylistic or privacy-driven edits.

BibTeX

@article{minh2026unsafe2safe,
  title={Unsafe2Safe: Controllable Image Anonymization for Downstream Utility},
  author={Minh Dinh Trong and SouYoung Jin},
  journal={CVPR},
  year={2026}
}