Hao Zhu1,2, Chaoyou Fu2, Qianyi Wu3, Wayne Wu3, Chen Qian3, Ran He2
1Anhui University | 2NLPR&CRIPAC, CASIA | 3SenseTime Research
Abstract: Recent studies have shown that the performance of forgery detection can be improved with diverse and challenging Deepfakes datasets. However, due to the lack of Deepfakes datasets with large variance in appearance, which can be hardly produced by recent identity swapping methods, the detection algorithm may fail in this situation. In this work, we provide a new identity swapping algorithm with large differences in appearance for face forgery detection. The appearance gaps mainly arise from the large discrepancies in illuminations and skin colors that widely exist in real-world scenarios. However, due to the difficulties of modeling the complex appearance mapping, it is challenging to transfer fine-grained appearances adaptively while preserving identity traits. This paper formulates appearance mapping as an optimal transport problem and proposes an Appearance Optimal Transport model (AOT) to formulate it in both latent and pixel space. Specifically, a relighting generator is designed to simulate the optimal transport plan. It is solved via minimizing the Wasserstein distance of the learned features in the latent space, enabling better performance and less computation than conventional optimization. To further refine the solution of the optimal transport plan, we develop a segmentation game to minimize the Wasserstein distance in the pixel space. A discriminator is introduced to distinguish the fake parts from a mix of real and fake image patches. Extensive experiments reveal that the superiority of our method when compared with state-of-the-art methods and the ability of our generated data to improve the performance of face forgery detection.
ANY RESOUCES PROVIDED HERE ARE NOT TO BE USED FOR MALICIOUS OR INAPPROPRIATE USE CASES.
The following datasets and the checkpoints of detectors can be downloaded from the BaiduYun (code: dzgc).
We currently provide 100 manipulated videos of FF++ by refining the results of DeepFaceLab and DeeperForensics-1.0 respectively.
The more manipulated videos are coming soon.
Binary detection accuracy of two video classification baselines: I3D and TSN on the hidden set provided by DeeperForensics-1.0.
-
We trained the baselines on four manipulated datasets of FF++ produced by DeepFakes, Face2Face, FaceSwap, and NeuralTextures (Green bars).
-
Then, we add 100 manipulated videos produced by our method to the training set. All detection accuracies are improved with the addition of our data. (Blue bars).
@inproceedings{aot2020neurips,
title={AOT: Appearance Optimal Transport Based Identity Swapping for Forgery Detection},
author={Zhu, Hao and Fu, Chaoyou and Wu, Qianyi and Wu, Wayne and Qian, Chen and He, Ran},
booktitle={Neural Information Processing Systems (NeurIPS)},
year={2020}
}
Coming soon.