Discrete Contrastive Diffusion for Cross-Modal Music and Image Generation
Ye Zhu Department of Computer Science Illinois Institute of Technology Chicago, IL 60616, USA、Yu Wu School of Computer Science Wuhan University Wuhan 430000, China \ANDKyle Olszewski, Jian Ren, Sergey Tulyakov Snap Inc. Santa Monica, CA 90405, USA、Yan Yan Department of Computer Science Illinois Institute of Technology Chicago, IL 60616, USA
Abstract
Diffusion probabilistic models (DPMs) have become a popular approach to conditional generation, due to their promising results and support for cross-modal synthesis. A key desideratum in conditional synthesis is to achieve high correspondence between the conditioning input and generated output. Most existing methods learn such relationships implicitly, by incorporating the prior into the variational lower bound. In this work, we take a different route—we explicitly enhance input-output connections by maximizing their mutual information. To this end, we introduce a Conditional Discrete Contrastive Diffusion (CDCD) loss and design two contrastive diffusion mechanisms to effectively incorporate it into the denoising process, combining the diffusion training and contrastive learning for the first time by connecting it with the conventional variational objectives. We demonstrate the efficacy of our approach in evaluations with diverse multimodal conditional synthesis tasks: dance-to-music generation, text-to-image synthesis, as well as class-conditioned image synthesis. On each, we enhance the input-output correspondence and achieve higher or competitive general synthesis quality. Furth
中文速览
扩散概率模型在跨模态条件生成任务中表现出色,但现有方法通常只是把条件信息隐式地塞进变分下界,导致输入条件与生成结果之间的对应关系有时会丢失。为此,作者提出通过最大化互信息来显式增强这种对应关系,设计了一种"条件离散对比扩散损失"(CDCD loss),并配套提出逐步并行扩散和逐样本辅助扩散两种机制,将对比学习首次融入扩散模型的去噪训练过程。在舞蹈转音乐、文本转图像、类别条件图像生成三类任务上的实验表明,该方法不仅提升了输入输出的对应质量和整体生成保真度,还让扩散模型收敛更快——在两个基准上所需扩散步数分别减少了约35%和40%,显著加快了推理速度。这项工作的意义在于,它提供了一种通用且有理论依据的方式,让扩散模型在条件生成时真正"听懂"输入,而不只是把条件当作一个附加的先验。
原文 arXiv:2206.07771;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2206.07771v2