EGSDE: Unpaired Image-to-Image Translation via Energy-Guided Stochastic Differential Equations
Min Zhao1, Fan Bao1, Chongxuan Li2,3 , Jun Zhu1∗ 1Dept. of Comp. Sci.、Tech., BNRist Center, THU-Bosch ML Center, Tsinghua University, China 2 Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China 3 Beijing Key Laboratory of Big Data Management and Analysis Methods , Beijing, China 4 Pazhou Laboratory (Huangpu), Guangzhou, China Correspondence to Chongxuan Li and Jun Zhu.
Abstract
Score-based diffusion models (SBDMs) have achieved the SOTA FID results in unpaired image-to-image translation (I2I). However, we notice that existing methods totally ignore the training data in the source domain, leading to sub-optimal solutions for unpaired I2I. To this end, we propose energy-guided stochastic differential equations (EGSDE) that employs an energy function pretrained on both the source and target domains to guide the inference process of a pretrained SDE for realistic and faithful unpaired I2I. Building upon two feature extractors, we carefully design the energy function such that it encourages the transferred image to preserve the domain-independent features and discard domain-specific ones. Further, we provide an alternative explanation of the EGSDE as a product of experts, where each of the three experts (corresponding to the SDE and two feature extractors) solely contributes to faithfulness or realism. Empirically, we compare EGSDE to a large family of baselines on three widely-adopted unpaired I2I tasks under four metrics. EGSDE not only consistently outperforms existing SBDMs-based methods in almost all settings but also achieves the SOTA realism results wit
中文速览
非配对图像翻译(unpaired image-to-image translation)长期面临一个被忽视的问题:现有基于扩散模型的方法在推理时完全不利用源域的训练数据,导致翻译结果在真实感和忠实度之间难以兼顾。研究者提出了能量引导随机微分方程(EGSDE),在预训练的目标域扩散模型推理过程中,引入一个同时在源域和目标域上预训练的能量函数作为引导信号——该能量函数由一个域特异性分类特征提取器和一个低通滤波器共同构成,前者促使翻译图像丢弃源域专有特征以提升真实感,后者督促模型保留姿态、颜色等域无关特征以提升忠实度。在 AFHQ 和 CelebA-HQ 数据集的多个翻译任务上,EGSDE 在 FID 等四项指标上全面超越已有扩散模型方法、并达到最优真实感水平,同时还支持通过超参数灵活调节真实感与忠实度之间的平衡,为利用双域信息改进生成式图像翻译提供了一条清晰可行的路径。
原文 arXiv:2207.06635;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2207.06635v5