How Much is Enough? A Study on Diffusion Times in Score-based Generative Models
Giulio Franzese EURECOM (France)、Simone Rossi EURECOM (France)、Lixuan Yang Huawei Technologies (France)、Alessandro Finamore Huawei Technologies (France)、Dario Rossi Huawei Technologies (France)、Maurizio Filippone EURECOM (France)、Pietro Michiardi EURECOM (France)
Abstract
Score-based diffusion models are a class of generative models whose dynamics is described by stochastic differential equations that map noise into data. While recent works have started to lay down a theoretical foundation for these models, an analytical understanding of the role of the diffusion time $T$ is still lacking. Current best practice advocates for a large $T$ to ensure that the forward dynamics brings the diffusion sufficiently close to a known and simple noise distribution; however, a smaller value of $T$ should be preferred for a better approximation of the score-matching objective and higher computational efficiency. Starting from a variational interpretation of diffusion models, in this work we quantify this trade-off, and suggest a new method to improve quality and efficiency of both training and sampling, by adopting smaller diffusion times. Indeed, we show how an auxiliary model can be used to bridge the gap between the ideal and the simulated forward dynamics, followed by a standard reverse diffusion process. Empirical results support our analysis; for image data, our method is competitive w.r.t. the state-of-the-art, according to standard sample quality metrics a
中文速览
扩散模型(score-based diffusion models)在生成高质量图像和音频方面表现出色,但始终存在一个被忽视的核心矛盾:扩散时间 T 越大,正向过程的终态越接近标准高斯噪声、反向过程的起点误差越小,但同时分数函数(score function)的学习难度会急剧上升、训练和采样的计算代价也随之飙升;T 越小则反过来。这篇论文从变分推断角度出发,对证据下界(ELBO)进行了新的分解,首次从理论上严格刻画了这一关于 T 的权衡关系,证明存在一个最优扩散时间使模型质量与计算效率同时达到最佳。在此基础上,作者提出了一种引入辅助模型的新方法:先用辅助模型将简单噪声分布"桥接"到正向过程实际终态,再接上标准反向扩散,从而在较小的 T 下消除初始条件不匹配带来的误差。在图像生成实验中,该方法在样本质量指标和对数似然两方面均与当前最优方法持平甚至更优,同时显著降低了训练与采样的计算开销。
原文 arXiv:2206.05173;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2206.05173v1