ConsistNet: Enforcing 3D Consistency for Multi-view Images Diffusion
Jiayu Yang1,212{}^{1,2}start_FLOATSUPERSCRIPT 1 , 2 end_FLOATSUPERSCRIPT, Ziang Cheng1,212{}^{1,2}start_FLOATSUPERSCRIPT 1 , 2 end_FLOATSUPERSCRIPT, Yunfei Duan11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT, Pan Ji11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT, Hongdong Li22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT 11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPTTencent, 22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPTAustralian National University {jiayu.yang, ziang.cheng,
Abstract
Given a single image of a 3D object, this paper proposes a novel method (named ConsistNet) that is able to generate multiple images of the same object, as if seen they are captured from different viewpoints, while the 3D (multi-view) consistencies among those multiple generated images are effectively exploited. Central to our method is a multi-view consistency block which enables information exchange across multiple single-view diffusion processes based on the underlying multi-view geometry principles. ConsistNet is an extension to the standard latent diffusion model, and consists of two sub-modules: (a) a view aggregation module that unprojects multi-view features into global 3D volumes and infer consistency, and (b) a ray aggregation module that samples and aggregate 3D consistent features back to each view to enforce consistency. Our approach departs from previous methods in multi-view image generation, in that it can be easily dropped-in pre-trained LDMs without requiring explicit pixel correspondences or depth prediction. Experiments show that our method effectively learns 3D consistency over a frozen Zero123 backbone and can generate 16 surrounding views of the object within
中文速览
从单张照片生成一个三维物体多角度、多视图图像时,现有扩散模型(如Zero123)缺乏有效机制来保证不同视角生成结果之间的三维几何一致性。ConsistNet提出了一个即插即用的多视图一致性模块:并行运行多个单视角潜扩散模型,通过"视角聚合"将各视角特征反投影到统一三维体素空间以推断全局一致性,再通过"射线聚合"将三维一致特征重新投影回各视角,以残差形式注入冻结的预训练网络,全程无需显式像素对应或深度预测。在Objaverse数据集上训练、Google Scanned Objects上测试,ConsistNet在PSNR、SSIM等多项指标上显著超越Zero123和SyncDreamer等方法,并可在单张A100 GPU上于40秒内生成16张环绕视图。这一工作意义在于提供了一种轻量、兼容现有预训练模型的三维一致多视图生成方案,可直接用于三维资产重建及VR/AR等下游应用。
原文 arXiv:2310.10343;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2310.10343v1