Code for the paper "MGI: Member vs Generated Inference" (ECCV 2026), which introduces the Member vs Generated Inference task (is a sample a training member of a generative model or an output generated by that model?) and Data Circuit Breaker (DCB), a three-stage method that combines signals from a generative model's autoencoder and its latent generator.
Abstract. As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data. We formalize this challenge as Member vs Generated Inference (MGI): given a sample and a target generative model, infer whether the sample is a true training member or a generated output of that model. Focusing on image generation, we show that existing membership inference methods systematically misclassify generated samples as training members, while attribution-based methods often misclassify true members as generated. This failure arises because both approaches rely on likelihood-related signals that are similarly elevated for training examples and for the model's own outputs. To address MGI, we propose Data Circuit Breaker (DCB), a three-stage method that combines complementary signals from a generative model's autoencoder and latent generator to distinguish training members from generated samples. Across multiple generative models, including image autoregressive and diffusion models, DCB consistently addresses the shortcomings of membership inference and attribution methods, remains effective even when models reproduce near-duplicates of training samples, and generalizes to challenging model derivative settings in which new models are trained on generated data.
- Installation
- Repository layout and paths
- Fine-tuned models
- Image sets
- Feature extraction
- Reproducing the tables
- Third-party code
- Citation
conda create -n mgi python=3.10 -y
conda activate mgi
pip install -r requirements.txtor, with uv:
uv venv --python 3.10
source .venv/bin/activate
uv pip install -r requirements.txtThe code was tested with Python 3.10, PyTorch 2.5.1 (CUDA 12.4), diffusers 0.36 and transformers 4.57 on NVIDIA A40 GPUs (46 GB). Every model fits on a single GPU with the default batch sizes.
Base weights are downloaded automatically from the Hugging Face Hub the first time a model is
used (VAR: FoundationVision/var, RAR: yucornetto/RAR and fun-research/TiTok, LlamaGen:
FoundationVision/LlamaGen, Stable Diffusion: CompVis/stable-diffusion-v1-4 and
Manojb/stable-diffusion-2-1-base, a mirror of stabilityai/stable-diffusion-2-1-base, which
is no longer available on the Hub). Set HF_HOME if you want the cache somewhere else.
.
├── configs
│ ├── paths.yaml # data_root, features_root, checkpoints_root
│ └── models/<model>.yaml # architecture, pre-processing, checkpoints and image sets per model
├── mgi # library
│ ├── models/ # model wrappers (VAR, RAR, LlamaGen, Stable Diffusion)
│ ├── config.py # config loading and path resolution
│ ├── data.py # image-set dataset and pre-processing
│ ├── extractor.py # feature extraction
│ ├── features.py # feature-file naming and loading
│ ├── scores.py # PIAR / CLiD, ICAS and DCB Stage-1 scores
│ ├── prada.py # PRADA scorer
│ ├── kde.py # DCB Stage-3 KDE test
│ ├── metrics.py # AUC and TPR@FPR
│ ├── tables.py # table assembly
│ └── cli.py # shared command-line arguments
├── scripts
│ ├── prepare_imagenet_subset.py # natural image sets for the IARs
│ ├── prepare_mscoco_subset.py # natural image-caption sets for Stable Diffusion
│ ├── generate_iar.py # sample from an IAR (optionally saving the latents)
│ ├── generate_sd.py # sample from Stable Diffusion
│ ├── generate_captions.py # BLIP-2 / LLaVA caption estimation (Table 7)
│ ├── finetune_iar.py # IAR model derivatives M2
│ ├── finetune_sd.py # Stable Diffusion fine-tuning (M1 and M2)
│ ├── finetune_encoder.py # refined IAR encoders (DCB Stage 1)
│ ├── extract_features.py # run one image set through one model
│ ├── extract_all.sh # every extraction needed for the tables
│ ├── table_direct_training.py # Tables 2 and 3
│ ├── table_model_derivative.py # Table 4
│ └── table_prompt_estimation.py # Table 7
├── prompts # MS-COCO captions used to generate the Stable Diffusion image sets
└── third_party # trimmed VAR / RAR / LlamaGen model definitions
configs/paths.yaml defines three roots (relative paths are resolved against the repository):
| Key | Default | Content |
|---|---|---|
data_root |
./data |
one sub-directory per image set (see Image sets) |
features_root |
./features |
extracted feature files |
checkpoints_root |
./checkpoints |
released checkpoints (see below) |
Every script also accepts --data-root, --features-root and --checkpoints-root.
The five models of the paper are var_30, rar_xxl, llamagen_c2i_xxl_384, sd1_4 and
sd2_1 (configs/models/). Each config lists the image sets of the paper under sets:, so
that the scripts can refer to a set by its role (natural_members, natural_nonmembers,
generated, generated_members, generated_nonmembers, m2_generated) instead of a path.
Two model variants are used throughout:
- M1 - the model that generated the data of the direct-training setting. For the IARs this is the public pre-trained checkpoint; for Stable Diffusion it is the model fine-tuned on 2,500 MS-COCO image-caption pairs (following CLiD), so that natural members and non-members are known.
- M2 - the model derivative: M1's latent generator fine-tuned on 5,000 images generated by M1.
Besides the public pre-trained weights (downloaded automatically), the experiments use fine-tuned
models that are produced with the scripts of this repository and stored under checkpoints_root
at the paths configured in configs/models/*.yaml:
Path under checkpoints_root |
Produced by | Purpose |
|---|---|---|
encoders/var_d30_vqvae_finetuned.pth |
scripts/finetune_encoder.py --model var_30 |
refined VAR encoder (DCB Stage 1) |
encoders/rar_maskgit_vqgan_encoder_delta.pth |
scripts/finetune_encoder.py --model rar_xxl |
refined RAR encoder (weight delta) |
encoders/llamagen_vq_ds16_c2i_encoder_finetuned.pt |
scripts/finetune_encoder.py --model llamagen_c2i_xxl_384 |
refined LlamaGen encoder |
var_30_M2/var_d30_M2.pth, rar_xxl_M2/rar_xxl_M2.pth, llamagen_c2i_xxl_384_M2/... |
scripts/finetune_iar.py |
IAR model derivatives M2 |
sd1_4_M1/, sd2_1_M1/ |
scripts/finetune_sd.py (base model, MS-COCO pairs) |
SD models whose training data is known |
sd1_4_M2/, sd2_1_M2/ |
scripts/finetune_sd.py (M1, images generated by M1) |
SD model derivatives M2 |
The commands and hyper-parameters are given in the next section. The refined encoders are
optional for the IARs (the paper's "gray-box" setting uses --encoder orig everywhere); the
diffusion models never use a refined encoder.
TODO: release the fine-tuned checkpoints, so that the tables can be reproduced without running the fine-tuning.
An image set is a directory in data_root with sorted image files (image_000000.png, ...
or MS-COCO *.jpg) plus either classes.npy (int64 ImageNet class of every image, for the
class-conditional IARs) or prompts.txt (one caption per image, for Stable Diffusion). Optional
prompts_blip2.txt / prompts_llava.txt files hold estimated captions. All sets are used in
sorted file-name order and the first 1,000 images of every set enter the tables.
# ImageNet: one training image per class (natural members N_M) and one validation image per class (N_N)
python scripts/prepare_imagenet_subset.py --imagenet-dir /path/to/imagenet/train --output data/imagenet_train
python scripts/prepare_imagenet_subset.py --imagenet-dir /path/to/imagenet/val --output data/imagenet_val
# MS-COCO (2014 validation images + captions): 2,500 pairs to fine-tune SD (N_M) and 2,500 disjoint pairs (N_N)
python scripts/prepare_mscoco_subset.py --coco-dir /path/to/coco2014 --output data/mscoco_train --num-images 2500 --seed 42
python scripts/prepare_mscoco_subset.py --coco-dir /path/to/coco2014 --output data/mscoco_val --num-images 2500 --seed 42 --exclude data/mscoco_trainFor each IAR (var_30, rar_xxl, llamagen_c2i_xxl_384; names below for VAR-d30) the paper uses
G_M = 5,000 images generated by M1 to fine-tune M2, G_N = 1,000 held-out M1 images (also the
generated set G of the direct-training setting) and G' = 1,000 images generated by M2. Labels
are sampled uniformly at random.
python scripts/generate_iar.py --model var_30 --output data/var30train_generated --num-samples 5000 --seed 1
python scripts/generate_iar.py --model var_30 --output data/var30test_generated --num-samples 1000 --seed 42
python scripts/finetune_iar.py --model var_30 --output checkpoints/var_30_M2 # 5 epochs, lr 1e-5, batch 4
python scripts/generate_iar.py --model var_30 --variant M2 --output data/var30M2test_generated --num-samples 1000 --seed 2finetune_iar.py trains only the transformer (the tokenizer stays frozen) with AdamW
(lr 1e-5, weight decay 0.03, cosine schedule with 100 warm-up steps, bf16 autocast) for 5 epochs
over generated_members, writing checkpoint-<epoch>.pt; point variants.M2 in the model
config to the final checkpoint (or rename it to the configured default, e.g.
checkpoints/var_30_M2/var_d30_M2.pth). The sampling parameters
of every model are listed under generation: in its config and follow the original works.
The encoder of each IAR is refined to invert its decoder on sampled latents
(L_inv = MSE(E(D(z)), z); the images themselves are not used). Sample latents with
--save-tokens and fine-tune (defaults reproduce the paper: VAR lr 5e-5 / batch 16 / 10 epochs;
RAR lr 5e-4 / batch 8 / 50 epochs; LlamaGen lr 1e-5 / batch 16 / 25 epochs; all with Adam and
StepLR(2, 0.9), trained on 50k sampled latents as in the example below):
python scripts/generate_iar.py --model var_30 --output data/var30enc_generated --num-samples 50000 --save-tokens --seed 500
python scripts/finetune_encoder.py --model var_30 --tokens data/var30enc_generated/tokens.ptThe output is written to the finetuned_encoder path of the model config and is picked up by
scripts/extract_features.py --encoder ft. For VAR and LlamaGen a few thousand sampled latents
already saturate the Stage-1 separation; RAR is more sensitive to the amount of training data
and needs the full 50k latents to reach the reported separation.
# M1: fine-tune the base model on the 2,500 MS-COCO pairs (50k steps, effective batch 4, lr 1e-5, EMA)
python scripts/finetune_sd.py --pretrained_model_name_or_path CompVis/stable-diffusion-v1-4 \
--dataset_name data/mscoco_train --output_dir checkpoints/sd1_4_M1 \
--resolution 512 --random_flip --train_batch_size 1 --gradient_accumulation_steps 4 \
--max_train_steps 50000 --learning_rate 1e-5 --lr_scheduler constant --lr_warmup_steps 0 \
--max_grad_norm 1 --use_ema --mixed_precision fp16 --gradient_checkpointing --checkpointing_steps 500000000
# G_M: 5,000 images from M1 (captions of COCO training images)
python scripts/generate_sd.py --model sd1_4 --variant M1 --prompts prompts/coco_prompts_train_first_5k.txt \
--num-samples 5000 --output data/sd14M2_train --seed 42
# G (= G_N): 2,500 images from M1 with the captions of the natural non-members
python scripts/generate_sd.py --model sd1_4 --variant M1 --prompts data/mscoco_val/prompts.txt \
--num-samples 2500 --output data/sd14ft_generated --seed 42
# M2: fine-tune M1 on G_M (20 epochs = 25k steps, batch 4, lr 1e-5, EMA)
python scripts/finetune_sd.py --pretrained_model_name_or_path checkpoints/sd1_4_M1 \
--dataset_name data/sd14M2_train --output_dir checkpoints/sd1_4_M2 \
--resolution 512 --random_flip --train_batch_size 4 --gradient_accumulation_steps 1 \
--num_train_epochs 20 --learning_rate 1e-5 --lr_scheduler constant --lr_warmup_steps 0 \
--max_grad_norm 1 --use_ema --mixed_precision fp16 --gradient_checkpointing --checkpointing_steps 500000000
# G': 1,000 images from M2 (captions of other COCO training images)
python scripts/generate_sd.py --model sd1_4 --variant M2 --prompts prompts/coco_prompts_train_second_5k.txt \
--num-samples 1000 --output data/sd14M2epoch20_generated --seed 42scripts/finetune_sd.py is the diffusers text-to-image fine-tuning example; it reads the
metadata.jsonl written by prepare_mscoco_subset.py / generate_sd.py. For SD 2.1 replace
sd1_4 by sd2_1, the base model by Manojb/stable-diffusion-2-1-base and the set names by
sd21M2_train, sd21M1.5_generated and sd21M2e20_generated (see configs/models/sd2_1.yaml).
Generation uses 50 DDIM-style inference steps with guidance scale 7.5; black (safety-filtered)
outputs are discarded and replaced.
for s in sd14M2_train sd14ft_generated sd14M2epoch20_generated; do
python scripts/generate_captions.py --caption-model blip2 --data-dir data/$s
python scripts/generate_captions.py --caption-model llava --data-dir data/$s --load-4bit
donewrites prompts_blip2.txt (BLIP-2 Flan-T5-XL) and prompts_llava.txt (LLaVA-1.6 Mistral-7B)
next to the images.
scripts/extract_features.py runs one image set through one model and stores, per image, the
conditional and unconditional per-token log-probabilities (per-timestep noise-prediction errors
for diffusion models), the VQ quantisation error and the double-reconstruction ratio of the
autoencoder:
# VAR-d30 scoring its natural members with the original encoder and the training-recipe resize
python scripts/extract_features.py --model var_30 --set natural_members --mid-reso
# ... with the refined encoder (DCB Stage 1)
python scripts/extract_features.py --model var_30 --set generated --encoder ft
# the model derivative M2 scoring its own generations
python scripts/extract_features.py --model rar_xxl --variant M2 --set m2_generated
# SD 1.4 scoring generated images with BLIP-2 captions at 20 noise levels (Table 7)
python scripts/extract_features.py --model sd1_4 --set generated_members --prompt-source blip2 --num-timesteps 20
# an arbitrary directory
python scripts/extract_features.py --model rar_xxl --data-dir data/rarxxlmemorig0.7_generated --num-samples 0Options: --set <role> or --data-dir <dir>; --variant M1|M2; --encoder orig|ft;
--mid-reso (VAR / LlamaGen: resize the shorter edge to 1.125x / 1.1x the input size before
the center crop, as in their training pipelines - used for all membership-inference scores);
--prompt-source default|blip2|llava; --num-timesteps (diffusion models, default 50);
--num-samples (default 1000, 0 = all); --seed (default 0, fixes the diffusion noise).
Feature files are written to
features_root/<model>[_M2]/<orig_enc|ft_enc>[_t<T>][_prompt_<source>]/<set>.pt, where <set>
is the directory name, with _mid appended (natural sets) or mid inserted before
_generated (generated sets) for --mid-reso.
scripts/extract_all.sh lists every extraction needed for the tables of the paper (about 40
GPU-hours on an A40 in total, most of it for the diffusion models; the commands are
independent and can be distributed over several GPUs).
All table scripts read the feature files, print a Markdown table (--latex for LaTeX rows) and
optionally write the raw numbers as JSON (--output). Values are TPR@1%FPR in percent unless
--metric auc or another --fpr is given; the best value per column is in bold. The rows are
the membership-inference baselines PIAR (IARs) / CLiD (diffusion models) and ICAS, the
attribution baseline PRADA, and DCB.
Score directions follow the assumptions of each family of methods: membership inference expects N_M > G > N_N (model derivative setting: G_M > G' > G_N > N_M > N_N, with N_M/N_N scored by M1 and all other pairs by M2), attribution (PRADA) expects G > N_M > N_N (G' > G_M > G_N > N_M > N_N), and DCB scores natural images above generated ones.
python scripts/table_direct_training.py --models rar_xxl var_30 llamagen_c2i_xxl_384 # Table 2
python scripts/table_direct_training.py --models sd1_4 sd2_1 # Table 3
python scripts/table_model_derivative.py --models var_30 rar_xxl sd1_4 sd2_1 # Table 4
python scripts/table_prompt_estimation.py --model sd1_4 --prompt-source blip2 # Table 7
python scripts/table_prompt_estimation.py --model sd1_4 --prompt-source llava
python scripts/table_prompt_estimation.py --model sd2_1 --prompt-source blip2
python scripts/table_prompt_estimation.py --model sd2_1 --prompt-source llava
python scripts/table_direct_training.py --models rar_xxl var_30 llamagen_c2i_xxl_384 --metric auc # appendix
python scripts/table_model_derivative.py --models var_30 rar_xxl sd1_4 sd2_1 --fpr 0.05 # appendixPRADA calibrates a small scorer with stochastic gradient descent; --prada-seed (default 0)
makes this deterministic. All other rows are deterministic functions of the feature files.
third_party/ contains trimmed copies of the model definitions of
VAR (MIT),
RAR / 1d-tokenizer (Apache-2.0) and
LlamaGen (MIT), extended with methods that expose
the pre- and post-quantisation latents of the tokenizers (third_party/README.md).
scripts/finetune_sd.py is derived from the diffusers text-to-image training example
(Apache-2.0). The feature extraction builds on
Privacy Attacks on Image AutoRegressive Models,
and the baselines follow the official descriptions of PIAR (Kowalczuk et al., 2025),
ICAS (Yu et al., 2025), CLiD (Zhai et al., 2024) and PRADA (Damm et al., 2026).
@inproceedings{zhao2026mgi,
title = {MGI: Member vs Generated Inference},
author = {Zhao, Bihe and Meintz, Michel and Xu, Juangui and Boenisch, Franziska and Dziedzic, Adam},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}