Abstract: Text-to-image diffusion models may memorize and reproduce training images. Recent mitigation methods try to suppress memorization by identifying and pruning localized weights. We show that this protection is fragile: after pruning, small changes to the text embeddings can recover retained training images. We call this attack DoRI. Our experiments indicate that memorization is not inherently local and that robust mitigation requires optimizing against adversarial retrieval triggers.
The easiest way to run the experiments is with Docker:
./docker_build.sh
./docker_run.sh --devices allThe container mounts the repository at /workspace. If your models, reference
images, or SSCD model are stored outside the repository, mount those folders as
well and pass their in-container paths to the scripts.
You can also use a local Python environment. Install a CUDA-compatible PyTorch build first, then install the remaining dependencies:
python -m pip install -r requirements.txtStable Diffusion models can be loaded by Hugging Face name, local path, or by one of the aliases below:
v1-4 -> CompVis/stable-diffusion-v1-4
v1-5 -> runwayml/stable-diffusion-v1-5
v2-0 -> stabilityai/stable-diffusion-2
v2-0-base -> stabilityai/stable-diffusion-2-base
If you keep local model folders under one root directory, set:
export DORI_MODEL_ROOT=/path/to/stable_diffusion_modelsFor SSCD metrics, download the SSCD TorchScript model separately and pass it
with --sscd_model_path:
wget -O /path/to/sscd_disc_mixup.torchscript.pt \
https://dl.fbaipublicfiles.com/sscd-copy-detection/sscd_disc_mixup.torchscript.ptThe prompt CSVs are included in prompts/. Provide local folders for
reference images, generated outputs, checkpoints, model files, SSCD, and any
datasets needed by the workflow you run.
Most scripts expect generated images to be named like:
img_0000_00.jpg
img_0000_01.jpg
img_0001_00.jpg
...
DoRI scripts sort reference images lexicographically, so the reference-image folder must be ordered consistently with the prompt CSV used for the run.
To download the memorized LAION reference images used by
prompts/memorized_laion_prompts.csv, run:
./download_memorized_reference_images.pyThis writes PNG files to reference_images/ using the expected ordering scheme:
0000_1030727993.png
0001_2026642560.png
...
Existing files are skipped. Any URLs that can no longer be downloaded are written
to reference_images/failed_downloads.csv.
Each script provides additional options via -h. The commands below use the
main-paper defaults where possible.
Set common paths:
export MODEL=v1-4
export OUT=outputs/v1_4
export REF=/path/to/local/reference_images
export SSCD=/path/to/sscd_disc_mixup.torchscript.pt
mkdir -p "${OUT}"First, identify memorization neurons. The script automatically performs the initial neuron selection and refinement:
python 3_detect_memorized_neurons.py \
--model "${MODEL}" \
--dataset prompts/memorized_laion_prompts.csv \
--output "${OUT}/memorization_statistics.csv"For MODEL=v1-4, this writes:
outputs/v1_4/memorization_statistics_v1_4.csv
Generate unmitigated images and NeMo-blocked images:
python 4_generate_images.py \
--model "${MODEL}" \
--result_file "${OUT}/memorization_statistics_v1_4.csv" \
--original_images \
--output "${OUT}/nemo_unblocked"
python 4_generate_images.py \
--model "${MODEL}" \
--result_file "${OUT}/memorization_statistics_v1_4.csv" \
--refined_neurons \
--output "${OUT}/nemo_blocked"Run DoRI against the NeMo-blocked model:
python 5_generate_images_adv_embeddings.py \
--model "${MODEL}" \
--result_file "${OUT}/memorization_statistics_v1_4.csv" \
--reference_images "${REF}" \
--output "${OUT}/nemo_dori"Optional: activation statistics for NeMo are already provided for SD v1.4 at
statistics/statistics_additional_laion_prompts_v1_4.pt. To recompute them:
python 1_compute_activations_statistics.py \
--model "${MODEL}" \
--prompts prompts/additional_laion_prompts.csv \
--output statistics/statistics_additional_laion_prompts.ptOptional: to recompute the pairwise SSIM threshold used during detection:
python 2_compute_pairwise_ssim.py \
--model "${MODEL}" \
--prompts prompts/additional_laion_prompts.csv \
--output outputs/pairwise_ssim_per_prompt.ptThe paper uses the default threshold 0.428.
First, collect Wanda input norms:
python wanda_01_get_input_norms.py \
--model "${MODEL}" \
--prompts prompts/memorized_laion_prompts.csv \
--output "${OUT}/wanda_input_norms"Generate Wanda-pruned images:
python wanda_02_generate_images.py \
--model "${MODEL}" \
--prompts prompts/memorized_laion_prompts.csv \
--input_norm_path "${OUT}/wanda_input_norms/input_norms.pkl" \
--sparsity 0.01 \
--timesteps_used 10 \
--output "${OUT}/wanda_blocked"Run DoRI against the Wanda-pruned model:
python wanda_04_attack.py \
--model "${MODEL}" \
--prompts prompts/memorized_laion_prompts.csv \
--input_norm_path "${OUT}/wanda_input_norms/input_norms.pkl" \
--memorized_images "${REF}" \
--sparsity 0.01 \
--timesteps_used 10 \
--output "${OUT}/wanda_dori"Use --memorization_type vm or --memorization_type tm for VM-only or TM-only
runs.
For memorized images, use the NeMo or Wanda DoRI commands above. For non-memorized controls, prepare local image/caption pairs outside git:
non_memorized_images/
0000000.jpg
0000000.txt
0000001.jpg
0000001.txt
Generate clean non-memorized images:
python 4_generate_images_non_mem.py \
--model "${MODEL}" \
--prompt_folder /path/to/local/non_memorized_images \
--num_images 500 \
--output "${OUT}/non_mem_clean"Run DoRI on non-memorized reference images:
python 5_generate_images_adv_embeddings_non_mem.py \
--model "${MODEL}" \
--image_folder /path/to/local/non_memorized_images \
--num_images 500 \
--output "${OUT}/non_mem_dori"The SD v2 experiments use the same NeMo, Wanda, and DoRI scripts with an SD2 model alias:
export MODEL=v2-0-base
export OUT=outputs/v2_0_baseFor NeMo detection on SD2, first compute SD2 activation statistics if they are not already available locally:
python 1_compute_activations_statistics.py \
--model "${MODEL}" \
--prompts prompts/additional_laion_prompts.csv \
--output statistics/statistics_additional_laion_prompts.ptThen run the NeMo, Wanda, DoRI, and metric commands above with SD2 prompts and reference images.
Prepare local training folders with image/caption pairs:
memorized_images/
0.jpg
0.txt
non_memorized_images/
0.jpg
0.txt
surrogate_images/
0.jpg
eval_images/
0.jpg
0.txt
Run adversarial fine-tuning:
python robust_finetuning.py \
--model "${MODEL}" \
--output_path "${OUT}/finetuning" \
--img_path_memorized /path/to/local/memorized_images \
--img_path_non_memorized /path/to/local/non_memorized_images \
--img_path_mitigated /path/to/local/surrogate_images \
--img_path_eval /path/to/local/eval_imagesGenerate images from a fine-tuned checkpoint:
python generate_images_after_fine_tuning.py \
--model "${MODEL}" \
--model_path "${OUT}/finetuning/checkpoints/checkpoint_epoch_5.pt" \
--prompts prompts/memorized_laion_prompts.csv \
--output "${OUT}/finetuned_generation"Run DoRI against the fine-tuned checkpoint:
python generate_images_after_fine_tuning.py \
--model "${MODEL}" \
--model_path "${OUT}/finetuning/checkpoints/checkpoint_epoch_5.pt" \
--prompts prompts/memorized_laion_prompts.csv \
--dori \
--memorized_image_path "${REF}" \
--output "${OUT}/finetuned_dori"For --dori, reference-image filenames must end with the corresponding prompt
index, for example sample_1030727993.jpg.
After generating images, compute the metrics in the metrics/ directory.
Measure similarity between generated images and original/reference images:
python metrics/compute_sscd_orig.py \
--folder "${OUT}/nemo_dori" \
--reference "${REF}" \
--prompts prompts/memorized_laion_prompts.csv \
--sscd_model_path "${SSCD}"Measure similarity between two generated folders:
python metrics/compute_sscd_gen.py \
--folder "${OUT}/nemo_dori" \
--reference "${OUT}/nemo_unblocked" \
--prompts prompts/memorized_laion_prompts.csv \
--sscd_model_path "${SSCD}"Measure sample diversity for each prompt:
python metrics/compute_diversity.py \
--folder "${OUT}/nemo_dori" \
--prompts prompts/memorized_laion_prompts.csv \
--sscd_model_path "${SSCD}"For FID, CLIP-FID, and KID, generate images for COCO prompts and evaluate them
with clean-fid. COCO captions are included in
prompts/coco2014_val_10000.csv; COCO images are not included. Generate the
30k COCO FID caption CSV locally from the official COCO annotations:
python download_coco_fid30k_captions.pyThis downloads annotations_trainval2014.zip, extracts
annotations/captions_val2014.json, and writes
MS-COCO_val2014_30k_captions.csv. The release tracks only the ordered image-id
manifest in prompts/coco_fid30k_image_ids.txt, not the generated caption CSV.
Prompt alignment is computed with CLIP:
python metrics/compute_prompt_alignment.py \
--folder "${OUT}/nemo_dori" \
--prompts prompts/memorized_laion_prompts.csvUseful checks:
git status --short
git ls-files | sortEach script also provides command-line help via -h.
Please cite the DoRI paper when using this code.
@inproceedings{kowalczuk2026dori,
title={Finding DoRI: Discovery of Retained Images in Diffusion Models},
author={Antoni Kowalczuk and Dominik Hintersdorf and Lukas Struppek and Kristian Kersting and Adam Dziedzic and Franziska Boenisch},
booktitle = {Proceedings of the 43rd International Conference on Machine Learning (ICML)},
year={2026}
}
