Discrete Contrastive Diffusion for Cross-Modal Music and Image Generation
Ye Zhu Affiliation: Department of Computer Science Affiliation: Illinois Institute of Technology Affiliation: Chicago, IL 60616, USA Email: Yu Wu Affiliation: School of Computer Science Affiliation: Wuhan University Affiliation: Wuhan 430000, China Email: Kyle Olszewski Jian Ren Sergey Tulyakov Affiliation: Snap Inc. Affiliation: Santa Monica, CA 90405, USA Email: Yan Yan Affiliation: Department of Computer Science Affiliation: Illinois Institute of Technology Affiliation: Chicago, IL 60616, USA Email:
Abstract
Diffusion probabilistic models (DPMs) have become a popular approach to conditional generation, due to their promising results and support for cross-modal synthesis. A key desideratum in conditional synthesis is to achieve high correspondence between the conditioning input and generated output. Most existing methods learn such relationships implicitly, by incorporating the prior into the variational lower bound. In this work, we take a different route—we explicitly enhance input-output connections by maximizing their mutual information. To this end, we introduce a Conditional Discrete Contrastive Diffusion (CDCD) loss and design two contrastive diffusion mechanisms to effectively incorporate it into the denoising process, combining the diffusion training and contrastive learning for the first time by connecting it with the conventional variational objectives. We demonstrate the efficacy of our approach in evaluations with diverse multimodal conditional synthesis tasks: dance-to-music generation, text-to-image synthesis, as well as class-conditioned image synthesis. On each, we enhance the input-output correspondence and achieve higher or competitive general synthesis quality. Furth
原文 arXiv:2206.07771;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2206.07771v2