Official code repository for the ICML 2026 accepted paper:
RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks
Paper link: https://arxiv.org/abs/2509.20924
RLCracker is an RL algorithm for evaluating the worst-case robustness of LLM text watermarks under adaptive paraphrasing attacks. It trains a detector-free rephrasing policy from a small number of prompt--watermarked response pairs and uses distributional rewards to move generated text away from watermark-induced behavior while preserving semantic fidelity.
This repository contains the code for reproducing the RLCracker training and evaluation pipeline.
Large language model watermarking is commonly evaluated against fixed prompts, standard paraphrasers, or non-adaptive rewriting attacks. RLCracker is designed as a stronger robustness stress test: it adaptively trains a rephraser to remove watermark signals without querying the watermark detector.
At a high level, RLCracker:
- trains from prompt--watermarked response pairs;
- does not require access to the watermark detector during training or evaluation;
- optimizes semantic preservation together with token-level distributional rewards;
- evaluates watermark robustness using evasion success rate, semantic similarity, and quality metrics;
- supports experiments across multiple watermarking schemes, model sizes, and text lengths.
The training code is built from Hugging Face Open-R1, and the watermarked data is generated with THU-BPM/MarkLLM. The MarkLLM/ directory in this repository stores the MarkLLM configs and datasets used by the experiments.
.
├── config_grpo_demo_reph.yaml # Example GRPO config
├── grpo_reph.py # Training entry point
├── GRPOTrainer_reph.py # Modified GRPO trainer and custom reward logic
├── evaluation.py # Rephrasing and watermark-detection evaluation
├── data_utils.py # Data preparation and result aggregation helpers
├── multi_thread.py # Multi-GPU batch launcher for training/evaluation
├── open_r1/ # Code adapted from Hugging Face Open-R1
├── watermarks/ # Watermark detection utilities
├── metrics/ # PSP, semantic similarity, and retrieval metrics
└── MarkLLM/ # MarkLLM configs and generated/used datasets
Create a Python 3.10+ environment and install the dependencies:
pip install -r requirements.txtSome runs require multi-GPU execution with accelerate, vllm, and trl. For gated Hugging Face models such as Llama, accept the model license and log in first:
huggingface-cli loginThe scripts use Hugging Face model IDs where possible. Common model IDs are:
Qwen/Qwen2.5-3B-InstructQwen/Qwen3-8Bmeta-llama/Meta-Llama-3-8B-Instructmeta-llama/Llama-3.1-8B-Instructsentence-transformers/all-MiniLM-L6-v2
Machine-specific absolute paths have been replaced by placeholders such as:
path/to/RLCracker/datasets/...
path/to/RLCracker/TRAINED_MODELS/...
Replace these placeholders with your local data, checkpoint, and result directories before running the scripts.
-
Generate watermarked text with MarkLLM using configs under
MarkLLM/config/and source datasets underMarkLLM/dataset/. -
Convert or filter the MarkLLM outputs into train/eval JSON files under a structure such as:
path/to/RLCracker/datasets/TRAIN_DATA/WM_GEN_Results_Filtered_short/<generator_model>/<watermark>/train_100.json
path/to/RLCracker/datasets/TRAIN_DATA/WM_GEN_Results_Filtered_short/<generator_model>/<watermark>/validation_20.json
path/to/RLCracker/datasets/TEST_DATA/WM_GEN_Results_Filtered_short/<generator_model>/Watermark_Test_<watermark>.json
Each training/evaluation record is expected to contain at least:
{
"question": "...",
"watermarked": "..."
}data_utils.py contains helper functions for preparing train/test splits, computing success rates, exporting TensorBoard scalars, and merging comparative Pass@K results.
This section describes how to run RLCracker training and evaluation, including single-run commands and the multi-GPU batch launcher in multi_thread.py.
grpo_reph.py loads the train/validation JSON files, wraps each watermarked field as a rephrasing prompt, and trains with the custom GRPOTrainer from GRPOTrainer_reph.py.
Example training command:
accelerate launch \
--config_file path/to/RLCracker/open_r1/accelerate_configs/zero2.yaml \
--num_processes 4 \
path/to/RLCracker/grpo_reph.py \
--config path/to/RLCracker/config_grpo_demo_reph.yaml \
--model_name_or_path Qwen/Qwen3-8B \
--max_completion_length 600 \
--min_completion_length 10 \
--per_device_eval_batch_size 6 \
--per_device_train_batch_size 6 \
--gradient_accumulation_steps 2 \
--vllm_tensor_parallel_size 4 \
--vllm_gpu_memory_utilization 0.08 \
--output_dir path/to/RLCracker/TRAINED_MODELS/Short/Meta-Llama-3.1-8B-Instruct/Qwen3_8B-EWD-Sem09KLRew01PPL5e7_dynaWeiSimAdv15 \
--temperature 0.7 \
--learning_rate 5e-7 \
--klreward_weight 0.9 \
--seman_weight 15 \
--ppl_weight 0.1 \
--train_data_path path/to/RLCracker/datasets/TRAIN_DATA/WM_GEN_Results_Filtered_short/Meta-Llama-3.1-8B-Instruct/EWD/train_100.json \
--test_data_path path/to/RLCracker/datasets/TRAIN_DATA/WM_GEN_Results_Filtered_short/Meta-Llama-3.1-8B-Instruct/EWD/validation_20.jsonIn this command:
--model_name_or_pathspecifies the rephraser model to train.--train_data_pathpoints to the 100-example training set.--test_data_pathpoints to the 20-example validation set.--klreward_weight,--seman_weight, and--ppl_weightcontrol the distributional reward, semantic reward, and perplexity penalty.--vllm_tensor_parallel_sizeshould usually match--num_processes.
Use evaluation.py to run a trained checkpoint as the rephraser and evaluate whether the rewritten outputs still trigger watermark detection.
Example evaluation command for short-form evaluation:
python path/to/RLCracker/evaluation.py \
--ckpt_path path/to/RLCracker/TRAINED_MODELS/Short/Meta-Llama-3.1-8B-Instruct/Qwen3_4B-KGW-Sem01KLRew01PPL1e6_dynaWeiSimAdv6/checkpoint-250 \
--data_path WM_GEN_Results_Filtered_short/Meta-Llama-3.1-8B-Instruct/KGW \
--run_num 1 \
--use_cuda True \
--is_short True \
--max_new_tokens 600 \
--min_new_tokens 100 \
--withSysPro True \
--WhetherThinking FalseExample evaluation command for long-form evaluation:
python path/to/RLCracker/evaluation.py \
--ckpt_path path/to/RLCracker/TRAINED_MODELS/Short/Meta-Llama-3.1-8B-Instruct/Qwen3_4B-KGW-Sem01KLRew01PPL1e6_dynaWeiSimAdv6/checkpoint-250 \
--data_path WM_GEN_Results_Filtered/Meta-Llama-3.1-8B-Instruct/KGW \
--run_num 1 \
--use_cuda True \
--is_short True \
--max_new_tokens 1600 \
--min_new_tokens 100 \
--withSysPro True \
--WhetherThinking FalseThe evaluation script writes JSONL files containing the source question, watermarked text, rephrased output, PSP score, sentence similarity score, and watermark detection results when detection is enabled for the selected watermark setting.
The repository also provides multi_thread.py, which launches multiple training or evaluation jobs across GPU groups.
Run:
python multi_thread.pyBy default, the script calls:
RunGRPO(node_num=4)
RunEval(node_num=1)The node_num argument controls how many GPUs are assigned to each job.
For example:
RunGRPO(node_num=4)launches each training job with 4 GPUs.RunEval(node_num=1)launches each evaluation job with 1 GPU.- If the machine has 8 GPUs and
node_num=4, the script creates two GPU groups:0,1,2,3and4,5,6,7. - If the machine has 8 GPUs and
node_num=1, the script creates eight GPU groups:0,1,2, ...,7.
Each GPU group runs one queued command at a time. The script sets:
CUDA_VISIBLE_DEVICES=<gpu_group>
VLLM_WORKER_MULTIPROC_METHOD=spawn
VLLM_DISABLE_COMPILE_CACHE=1for each job.
RunGRPO(node_num) creates a queue of training commands and assigns them to GPU groups.
The default training loop in multi_thread.py is:
for algorithm in ['EWD']: # 'SWEET','PF'
for model_name in ['Meta-Llama-3.1-8B-Instruct']:
...The corresponding training command template is:
accelerate launch \
--main_process_port <port> \
--config_file path/to/RLCracker/open_r1/accelerate_configs/zero2.yaml \
--num_processes <node_num> \
path/to/RLCracker/grpo_reph.py \
--config path/to/RLCracker/config_grpo_demo_reph.yaml \
--model_name_or_path Qwen/Qwen3-8B \
--max_completion_length 600 \
--min_completion_length 10 \
--per_device_eval_batch_size <24/node_num> \
--per_device_train_batch_size <24/node_num> \
--gradient_accumulation_steps 2 \
--vllm_tensor_parallel_size <node_num> \
--vllm_gpu_memory_utilization 0.08 \
--output_dir path/to/RLCracker/TRAINED_MODELS/Short/<generator_model>/Qwen3_8B-<watermark>-Sem09KLRew01PPL5e7_dynaWeiSimAdv15 \
--temperature 0.7 \
--learning_rate 5e-7 \
--klreward_weight 0.9 \
--seman_weight 15 \
--ppl_weight 0.1 \
--train_data_path path/to/RLCracker/datasets/TRAIN_DATA/WM_GEN_Results_Filtered_short/<generator_model>/<watermark>/train_100.json \
--test_data_path path/to/RLCracker/datasets/TRAIN_DATA/WM_GEN_Results_Filtered_short/<generator_model>/<watermark>/validation_20.jsonTo train on more watermarking schemes, edit:
for algorithm in ['EWD']:to, for example:
for algorithm in ['EWD', 'SWEET', 'PF']:To train on another generator model, edit:
for model_name in ['Meta-Llama-3.1-8B-Instruct']:To change the attacker model, edit:
--model_name_or_path Qwen/Qwen3-8Band update the output_dir name accordingly.
Each generated command is saved to:
<output_dir>/command.txt
which makes it easier to reproduce or debug a run.
RunEval(node_num) creates a queue of evaluation commands and assigns them to GPU groups.
The default evaluation loop in multi_thread.py evaluates checkpoints:
for ckpt_step in [25,50,75,100,125,150,175,200,225,250]:for the setting:
mod = 'Qwen3_4B'
algo = 'KGW'
model_name = 'Meta-Llama-3.1-8B-Instruct'
ckpt_name = 'Sem01KLRew01PPL1e6_dynaWeiSimAdv6'For each checkpoint, the script runs three evaluation commands:
python path/to/RLCracker/evaluation.py \
--ckpt_path path/to/RLCracker/TRAINED_MODELS/Short/Meta-Llama-3.1-8B-Instruct/Qwen3_4B-KGW-Sem01KLRew01PPL1e6_dynaWeiSimAdv6/checkpoint-<step> \
--data_path WM_GEN_Results_Filtered_short/Meta-Llama-3.1-8B-Instruct/KGW \
--run_num 1 \
--use_cuda True \
--is_short True \
--max_new_tokens 600 \
--min_new_tokens 100 \
--withSysPro True \
--WhetherThinking Falsepython path/to/RLCracker/evaluation.py \
--ckpt_path path/to/RLCracker/TRAINED_MODELS/Short/Meta-Llama-3.1-8B-Instruct/Qwen3_4B-KGW-Sem01KLRew01PPL1e6_dynaWeiSimAdv6/checkpoint-<step> \
--data_path WM_GEN_Results_Filtered/Meta-Llama-3.1-8B-Instruct/KGW \
--run_num 1 \
--use_cuda True \
--is_short True \
--max_new_tokens 1600 \
--min_new_tokens 100 \
--withSysPro False \
--WhetherThinking Falsepython path/to/RLCracker/evaluation.py \
--ckpt_path path/to/RLCracker/TRAINED_MODELS/Short/Meta-Llama-3.1-8B-Instruct/Qwen3_4B-KGW-Sem01KLRew01PPL1e6_dynaWeiSimAdv6/checkpoint-<step> \
--data_path WM_GEN_Results_Filtered/Meta-Llama-3.1-8B-Instruct/KGW \
--run_num 1 \
--use_cuda True \
--is_short True \
--max_new_tokens 1600 \
--min_new_tokens 100 \
--withSysPro True \
--WhetherThinking FalseTo evaluate a different watermark, change:
for algo in ['KGW']:To evaluate a different model or checkpoint naming pattern, change:
for mod in ['Qwen3_4B']:
for ckpt_name in ['Sem01KLRew01PPL1e6_dynaWeiSimAdv6']:To evaluate only selected checkpoints, edit:
for ckpt_step in [25,50,75,100,125,150,175,200,225,250]:A typical workflow is:
- Generate watermarked data with MarkLLM.
- Prepare
train_100.jsonandvalidation_20.json. - Run single training for debugging.
- Run
RunGRPO(node_num=4)for batch training. - Run one evaluation command manually to verify paths.
- Run
RunEval(node_num=1)for checkpoint evaluation. - Aggregate results with
data_utils.py.
The main evaluation metric is evasion success rate, which measures the fraction of rewritten outputs that both evade watermark detection and preserve semantic similarity.
The repository also supports additional metrics used in the paper, including:
- PSP semantic similarity;
- sentence-level similarity;
- watermark detection result before and after rewriting;
- removal rate without semantic filtering;
- quality-related statistics used for result aggregation.
Metric utilities are implemented under metrics/, and aggregation helpers are provided in data_utils.py.
- Generate watermarked data with MarkLLM.
- Prepare
TRAIN_DATAandTEST_DATAJSON files withquestionandwatermarkedfields. - Train the rephraser with
grpo_reph.py. - Evaluate checkpoints with
evaluation.py. - Aggregate success rates or Pass@K results with
data_utils.py.
RLCracker is released for research on watermark robustness, security evaluation, and pre-deployment stress testing. The code is intended to help watermark designers identify vulnerabilities in existing schemes and develop stronger defenses.
Because adaptive watermark-removal methods may be misused to evade provenance mechanisms, users should apply this repository only in authorized research or evaluation settings and follow the policies of the models, datasets, and watermarking systems they use.
open_r1/is adapted from Hugging Face Open-R1.MarkLLM/stores the MarkLLM configs and data used by this project.- PSP metric files are expected under
metrics/p_sp_utils/psp/; placemodel.para.lc.100.ptthere if it is not included. - The batch launcher in
multi_thread.pyis an example template. Adjust model names, watermark names, checkpoint steps, and GPU counts for your run. - When using
multi_thread.py, verify thatnode_num, available GPU count,vllm_tensor_parallel_size, and batch sizes are compatible.
If you find this repository useful, please cite:
@misc{huang2025rlcrackerexposingvulnerabilityllm,
title={RLCracker: Exposing the Vulnerability of LLM Watermarks with Adaptive RL Attacks},
author={Hanbo Huang and Yiran Zhang and Hao Zheng and Xuan Gong and Yihan Li and Lin Liu and Shiyu Liang},
year={2025},
eprint={2509.20924},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2509.20924},
}