Skip to content

Fix broken PhotoMaker v1 path: version-aware num_tokens + pass pm_version="v1" where the v1 checkpoint is loaded - #231

Open
Linxiushen wants to merge 1 commit into
TencentARC:mainfrom
Linxiushen:fix-photomaker-v1-compat
Open

Linxiushen wants to merge 1 commit into
TencentARC:mainfrom
Linxiushen:fix-photomaker-v1-compat

Conversation

@Linxiushen

Copy link
Copy Markdown

Problem

load_photomaker_adapter() defaults pm_version="v2" and then hardcodes self.num_tokens = 2. But the v1 flow the README still documents, both demo notebooks, and predict.py all download photomaker-v1.bin and call the loader without pm_version, so a v1 checkpoint is loaded down the v2 code path. Two failures follow:

  1. Setup fails on the state dict (this is what issues RuntimeError: Error(s) in loading state_dict for PhotoMakerIDEncoder_CLIPInsightfaceExtendtoken #194, The Notebook demo can't work #168, Error: photomaker_style_demo in qformer_perceiver.perceiver_resampler.layers.3.1.3.weigh #175 and error in running demo note book #193 report).
  2. Even past that, the v1 ID encoder produces one embedding per ID image while the pipeline reserved two class-token slots, so PhotoMakerIDEncoder.forward raises:
RuntimeError: Sizes of tensors must match except in dimension 1.
Expected size 2 but got size 1 for tensor number 1 in the list.

So fixing only the version argument just moves the crash later; both parts are needed, which is why this is one PR rather than two.

Reachability, stated plainly: the README quick start (lines 108-160) and the two notebooks it links are the affected paths, and copying them fails immediately. I am not claiming the hosted Replicate endpoint is down: its image was built in 2024-01 from v1-era code and is frozen, and cog.yaml pins diffusers==0.25.0, which cannot even import today's photomaker/pipeline.py (diffusers.callbacks only exists from 0.28). The predict.py line is still worth fixing, but not because production is broken.

Fix

  • self.num_tokens = 2 if pm_version == "v2" else 1, which restores the v1 semantics the code had before the v2 commit ([class_token] * num_id_images).
  • Pass pm_version="v1" at the four call sites that load photomaker-v1.bin: the README snippet, both notebooks, predict.py.
  • The same hardcoded num_tokens is fixed in pipeline_controlnet.py and pipeline_t2i_adapter.py for consistency. Stated plainly: those two have no v1 caller in the repo today (both inference_pmv2_* scripts fetch photomaker-v2.bin and use the v2 default), so they are latent, not currently broken.

The loader signature and its pm_version="v2" default are unchanged, so every v2 caller (the four inference_pmv2*.py scripts and gradio_demo/app_v2.py) is untouched.

Verification

zsh verify.sh in my scratch dir reproduces the whole thing; the important parts:

  • v2 unchanged: 5 prompt templates x num_id_images in {1,2,3,4,5,8} (30 combinations), comparing clean_input_ids and class_tokens_mask between the pristine and patched trees: all identical. End-to-end forward through the real, unmodified PhotoMakerIDEncoder_CLIPInsightfaceExtendtoken with a fixed seed: max absolute difference 0.0, torch.equal true.
  • v1 fixed: with num_id_images of 1, 2, 4 and 8, the pristine tree raises the RuntimeError above every time; patched, all four return out=(1, 77, 2048).
  • Notebook edits are surgical insertions into the original JSON (not a json.dump rewrite), so the diff is one added and one changed line each and both still parse. README.md is CRLF and the patch preserves that.
  • ruff with the repo's pyproject.toml settings reports the same 221 findings before and after, with an identical per-rule breakdown; no new diagnostics.

No test file: the repo has none, requirements.txt has no pytest, and there is no CI, so testing this end to end would mean mocking the whole diffusers pipeline or downloading multi-GB weights. The differential script above covers it instead.

Note for whoever picks this up: open PR #200 touches photomaker/pipeline.py heavily, so if that lands first this needs a trivial textual rebase.


This fix was developed with AI assistance (Claude); the change was reviewed and verified locally before submission.

…sion="v1" where the v1 checkpoint is loaded

load_photomaker_adapter() defaults pm_version to "v2" and hardcodes
self.num_tokens = 2, but the README quick start, both demo notebooks and
predict.py all download photomaker-v1.bin and call it without
pm_version. So the documented v1 flow loads a v1 checkpoint under the v2
code path: setup fails on the state dict, and once that is past, the v1
encoder produces one embedding per ID image while the pipeline reserved
two class-token slots, so PhotoMakerIDEncoder.forward raises

    RuntimeError: Sizes of tensors must match except in dimension 1.
    Expected size 2 but got size 1 for tensor number 1 in the list.

Make num_tokens version-aware (`2 if pm_version == "v2" else 1`, which
restores the pre-v2 v1 semantics) and pass pm_version="v1" at the four
call sites that load the v1 checkpoint: the README snippet, both
notebooks and predict.py.

The same hardcoded num_tokens is fixed in pipeline_controlnet.py and
pipeline_t2i_adapter.py for consistency; those two have no v1 caller in
the repo today, so they are latent rather than currently broken.

v2 behaviour is unchanged: the expression is identically 2 for "v2".
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant