Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion Priors
Yukang Lin, Haonan Han, Chaoqun Gong, Zunnan Xu, Yachao Zhang, Xiu Li Tsinghua Shenzhen International Graduate School, Tsinghua University {linyk23, hhn22, gcq22, {yachaozhang,
Abstract
Reconstructing 3D objects from a single image guided by pretrained diffusion models has demonstrated promising outcomes. However, due to utilizing the case-agnostic rigid strategy, their generalization ability to arbitrary cases and the 3D consistency of reconstruction are still poor. In this work, we propose Consistent123, a case-aware two-stage method for highly consistent 3D asset reconstruction from one image with both 2D and 3D diffusion priors. In the first stage, Consistent123 utilizes only 3D structural priors for sufficient geometry exploitation, with a CLIP-based case-aware adaptive detection mechanism embedded within this process. In the second stage, 2D texture priors are introduced and progressively take on a dominant guiding role, delicately sculpting the details of the 3D model. Consistent123 aligns more closely with the evolving trends in guidance requirements, adaptively providing adequate 3D geometric initialization and suitable 2D texture refinement for different objects. Consistent123 can obtain highly 3D-consistent reconstruction and exhibits strong generalization ability across various objects. Qualitative and quantitative experiments show that our method sign
中文速览
从单张图片重建出高质量、多视角一致的3D模型,一直是计算机视觉领域的难题,现有方法往往对所有物体套用同一套固定优化策略,导致几何结构混乱或出现"多面孔"等纹理失真问题。为此,研究者提出了Consistent123,一种感知具体案例(case-aware)的两阶段方法:第一阶段只用3D结构先验(3D diffusion prior)来稳定地恢复物体几何形状,并通过一个基于CLIP的自适应检测机制判断形状是否已充分重建;一旦检测通过,第二阶段便逐步引入2D纹理先验(2D diffusion prior),以动态权重调度的方式精细雕琢颜色和细节,同时不断降低3D先验的比重。在RealFusion15和自建C10数据集上的定量与定性实验均表明,Consistent123在几何一致性、纹理质量和泛化能力上显著超越现有最优方法,且无需人工调节先验比例。这项工作的价值在于,它为单图生成3D资产提供了一条兼顾结构忠实度与纹理细节的自适应路径,有望大幅降低3D内容创作的门槛。
原文 arXiv:2309.17261;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2309.17261v2