Skip to content

About

[ICML 2026] Source code for our paper "Finding DoRI: Discovery of Retained Images in Diffusion Models"

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Finding DoRI: Discovery of Retained Images in Diffusion Models (ICML 2026)

Concept

Abstract: Text-to-image diffusion models may memorize and reproduce training images. Recent mitigation methods try to suppress memorization by identifying and pruning localized weights. We show that this protection is fragile: after pruning, small changes to the text embeddings can recover retained training images. We call this attack DoRI. Our experiments indicate that memorization is not inherently local and that robust mitigation requires optimizing against adversarial retrieval triggers.

Setup

The easiest way to run the experiments is with Docker:

./docker_build.sh
./docker_run.sh --devices all

The container mounts the repository at /workspace. If your models, reference images, or SSCD model are stored outside the repository, mount those folders as well and pass their in-container paths to the scripts.

You can also use a local Python environment. Install a CUDA-compatible PyTorch build first, then install the remaining dependencies:

python -m pip install -r requirements.txt

Stable Diffusion models can be loaded by Hugging Face name, local path, or by one of the aliases below:

v1-4       -> CompVis/stable-diffusion-v1-4
v1-5       -> runwayml/stable-diffusion-v1-5
v2-0       -> stabilityai/stable-diffusion-2
v2-0-base  -> stabilityai/stable-diffusion-2-base

If you keep local model folders under one root directory, set:

export DORI_MODEL_ROOT=/path/to/stable_diffusion_models

For SSCD metrics, download the SSCD TorchScript model separately and pass it with --sscd_model_path:

wget -O /path/to/sscd_disc_mixup.torchscript.pt \
  https://dl.fbaipublicfiles.com/sscd-copy-detection/sscd_disc_mixup.torchscript.pt

Data

The prompt CSVs are included in prompts/. Provide local folders for reference images, generated outputs, checkpoints, model files, SSCD, and any datasets needed by the workflow you run.

Most scripts expect generated images to be named like:

img_0000_00.jpg
img_0000_01.jpg
img_0001_00.jpg
...

DoRI scripts sort reference images lexicographically, so the reference-image folder must be ordered consistently with the prompt CSV used for the run.

To download the memorized LAION reference images used by prompts/memorized_laion_prompts.csv, run:

./download_memorized_reference_images.py

This writes PNG files to reference_images/ using the expected ordering scheme:

0000_1030727993.png
0001_2026642560.png
...

Existing files are skipped. Any URLs that can no longer be downloaded are written to reference_images/failed_downloads.csv.

Reproducing the Main Experiments

Each script provides additional options via -h. The commands below use the main-paper defaults where possible.

Set common paths:

export MODEL=v1-4
export OUT=outputs/v1_4
export REF=/path/to/local/reference_images
export SSCD=/path/to/sscd_disc_mixup.torchscript.pt
mkdir -p "${OUT}"

1. NeMo Mitigation and DoRI Attack

First, identify memorization neurons. The script automatically performs the initial neuron selection and refinement:

python 3_detect_memorized_neurons.py \
  --model "${MODEL}" \
  --dataset prompts/memorized_laion_prompts.csv \
  --output "${OUT}/memorization_statistics.csv"

For MODEL=v1-4, this writes:

outputs/v1_4/memorization_statistics_v1_4.csv

Generate unmitigated images and NeMo-blocked images:

python 4_generate_images.py \
  --model "${MODEL}" \
  --result_file "${OUT}/memorization_statistics_v1_4.csv" \
  --original_images \
  --output "${OUT}/nemo_unblocked"

python 4_generate_images.py \
  --model "${MODEL}" \
  --result_file "${OUT}/memorization_statistics_v1_4.csv" \
  --refined_neurons \
  --output "${OUT}/nemo_blocked"

Run DoRI against the NeMo-blocked model:

python 5_generate_images_adv_embeddings.py \
  --model "${MODEL}" \
  --result_file "${OUT}/memorization_statistics_v1_4.csv" \
  --reference_images "${REF}" \
  --output "${OUT}/nemo_dori"

Optional: activation statistics for NeMo are already provided for SD v1.4 at statistics/statistics_additional_laion_prompts_v1_4.pt. To recompute them:

python 1_compute_activations_statistics.py \
  --model "${MODEL}" \
  --prompts prompts/additional_laion_prompts.csv \
  --output statistics/statistics_additional_laion_prompts.pt

Optional: to recompute the pairwise SSIM threshold used during detection:

python 2_compute_pairwise_ssim.py \
  --model "${MODEL}" \
  --prompts prompts/additional_laion_prompts.csv \
  --output outputs/pairwise_ssim_per_prompt.pt

The paper uses the default threshold 0.428.

2. Wanda Mitigation and DoRI Attack

First, collect Wanda input norms:

python wanda_01_get_input_norms.py \
  --model "${MODEL}" \
  --prompts prompts/memorized_laion_prompts.csv \
  --output "${OUT}/wanda_input_norms"

Generate Wanda-pruned images:

python wanda_02_generate_images.py \
  --model "${MODEL}" \
  --prompts prompts/memorized_laion_prompts.csv \
  --input_norm_path "${OUT}/wanda_input_norms/input_norms.pkl" \
  --sparsity 0.01 \
  --timesteps_used 10 \
  --output "${OUT}/wanda_blocked"

Run DoRI against the Wanda-pruned model:

python wanda_04_attack.py \
  --model "${MODEL}" \
  --prompts prompts/memorized_laion_prompts.csv \
  --input_norm_path "${OUT}/wanda_input_norms/input_norms.pkl" \
  --memorized_images "${REF}" \
  --sparsity 0.01 \
  --timesteps_used 10 \
  --output "${OUT}/wanda_dori"

Use --memorization_type vm or --memorization_type tm for VM-only or TM-only runs.

3. Memorized vs. Non-Memorized DoRI Controls

For memorized images, use the NeMo or Wanda DoRI commands above. For non-memorized controls, prepare local image/caption pairs outside git:

non_memorized_images/
  0000000.jpg
  0000000.txt
  0000001.jpg
  0000001.txt

Generate clean non-memorized images:

python 4_generate_images_non_mem.py \
  --model "${MODEL}" \
  --prompt_folder /path/to/local/non_memorized_images \
  --num_images 500 \
  --output "${OUT}/non_mem_clean"

Run DoRI on non-memorized reference images:

python 5_generate_images_adv_embeddings_non_mem.py \
  --model "${MODEL}" \
  --image_folder /path/to/local/non_memorized_images \
  --num_images 500 \
  --output "${OUT}/non_mem_dori"

4. Stable Diffusion v2 Experiments

The SD v2 experiments use the same NeMo, Wanda, and DoRI scripts with an SD2 model alias:

export MODEL=v2-0-base
export OUT=outputs/v2_0_base

For NeMo detection on SD2, first compute SD2 activation statistics if they are not already available locally:

python 1_compute_activations_statistics.py \
  --model "${MODEL}" \
  --prompts prompts/additional_laion_prompts.csv \
  --output statistics/statistics_additional_laion_prompts.pt

Then run the NeMo, Wanda, DoRI, and metric commands above with SD2 prompts and reference images.

5. Adversarial Fine-Tuning Mitigation

Prepare local training folders with image/caption pairs:

memorized_images/
  0.jpg
  0.txt
non_memorized_images/
  0.jpg
  0.txt
surrogate_images/
  0.jpg
eval_images/
  0.jpg
  0.txt

Run adversarial fine-tuning:

python robust_finetuning.py \
  --model "${MODEL}" \
  --output_path "${OUT}/finetuning" \
  --img_path_memorized /path/to/local/memorized_images \
  --img_path_non_memorized /path/to/local/non_memorized_images \
  --img_path_mitigated /path/to/local/surrogate_images \
  --img_path_eval /path/to/local/eval_images

Generate images from a fine-tuned checkpoint:

python generate_images_after_fine_tuning.py \
  --model "${MODEL}" \
  --model_path "${OUT}/finetuning/checkpoints/checkpoint_epoch_5.pt" \
  --prompts prompts/memorized_laion_prompts.csv \
  --output "${OUT}/finetuned_generation"

Run DoRI against the fine-tuned checkpoint:

python generate_images_after_fine_tuning.py \
  --model "${MODEL}" \
  --model_path "${OUT}/finetuning/checkpoints/checkpoint_epoch_5.pt" \
  --prompts prompts/memorized_laion_prompts.csv \
  --dori \
  --memorized_image_path "${REF}" \
  --output "${OUT}/finetuned_dori"

For --dori, reference-image filenames must end with the corresponding prompt index, for example sample_1030727993.jpg.

Evaluation Metrics

After generating images, compute the metrics in the metrics/ directory.

Memorization

Measure similarity between generated images and original/reference images:

python metrics/compute_sscd_orig.py \
  --folder "${OUT}/nemo_dori" \
  --reference "${REF}" \
  --prompts prompts/memorized_laion_prompts.csv \
  --sscd_model_path "${SSCD}"

Measure similarity between two generated folders:

python metrics/compute_sscd_gen.py \
  --folder "${OUT}/nemo_dori" \
  --reference "${OUT}/nemo_unblocked" \
  --prompts prompts/memorized_laion_prompts.csv \
  --sscd_model_path "${SSCD}"

Diversity

Measure sample diversity for each prompt:

python metrics/compute_diversity.py \
  --folder "${OUT}/nemo_dori" \
  --prompts prompts/memorized_laion_prompts.csv \
  --sscd_model_path "${SSCD}"

Quality and Prompt Alignment

For FID, CLIP-FID, and KID, generate images for COCO prompts and evaluate them with clean-fid. COCO captions are included in prompts/coco2014_val_10000.csv; COCO images are not included. Generate the 30k COCO FID caption CSV locally from the official COCO annotations:

python download_coco_fid30k_captions.py

This downloads annotations_trainval2014.zip, extracts annotations/captions_val2014.json, and writes MS-COCO_val2014_30k_captions.csv. The release tracks only the ordered image-id manifest in prompts/coco_fid30k_image_ids.txt, not the generated caption CSV.

Prompt alignment is computed with CLIP:

python metrics/compute_prompt_alignment.py \
  --folder "${OUT}/nemo_dori" \
  --prompts prompts/memorized_laion_prompts.csv

Validation

Useful checks:

git status --short
git ls-files | sort

Each script also provides command-line help via -h.

Citation

Please cite the DoRI paper when using this code.

@inproceedings{kowalczuk2026dori,
    title={Finding DoRI: Discovery of Retained Images in Diffusion Models},
    author={Antoni Kowalczuk and Dominik Hintersdorf and Lukas Struppek and Kristian Kersting and Adam Dziedzic and Franziska Boenisch},
    booktitle = {Proceedings of the 43rd International Conference on Machine Learning (ICML)},
    year={2026}
}

About

[ICML 2026] Source code for our paper "Finding DoRI: Discovery of Retained Images in Diffusion Models"

Resources

Stars

1 star

Watchers

0 watching

Forks

Contributors

Languages