TOSS: High-quality Text-guided NOvel View Synthesis from a Single Image
Yukai Shi1,3 Jianan Wang3111 He Cao2,3111 ††{}^{\,\,{\dagger}} Boshi Tang1,3 Xianbiao Qi3 Tianyu Yang3 Yukun Huang3 Shilong Liu1,3 Lei Zhang 3 Heung-Yeung Shum 1,3 1 Tsinghua University 2 Hong Kong University of Science and Technology 3 International Digital Economy Academy (IDEA) Equal contribution.Work done during an internship at IDEA.
Abstract
In this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image. While Zero123 has demonstrated impressive zero-shot open-set NVS capability, it treats NVS as a pure image-to-image translation problem. This approach suffers from the challengingly under-constrained nature of single-view NVS: the process lacks means of explicit user control and often results in implausible NVS generations. To address this limitation, TOSS uses text as high-level semantic information to constrain the NVS solution space. TOSS fine-tunes text-to-image Stable Diffusion pre-trained on large-scale text-image pairs and introduces modules specifically tailored to image and camera pose conditioning, as well as dedicated training for pose correctness and preservation of fine details. Comprehensive experiments are conducted with results showing that our proposed TOSS outperforms Zero123 with more plausible, controllable and multiview-consistent NVS results. We further support these results with comprehensive ablations that underscore the effectiveness and potential of the introduced semantic guidance and architecture design. See https://toss3d.github.
中文速览
单张RGB照片合成新视角(Novel View Synthesis, NVS)本质上是一个严重欠约束的问题——仅凭一张图根本无法唯一确定物体看不见的那一面长什么样,导致现有方法(如Zero123)常常生成语义上荒谬的结果,比如鲨鱼长出两条尾巴,且用户对生成内容毫无控制权。本文提出TOSS,核心思路是引入文本描述作为高层语义约束,同时设计专门的图像特征与相机位姿条件注入模块,并在Objaverse三维数据集上对Stable Diffusion进行微调,还用两个专家扩散模型分别保证视角准确性和细节保真度。实验结果表明,TOSS在视角合理性、用户可控性和多视角一致性上均明显优于Zero123,同时支持文本反演等进一步优化手段。这项工作的意义在于,它为单图三维理解和三维内容生成提供了一个更可靠、更灵活的基础能力,有望成为未来三维生成流水线的重要组成部分。
原文 arXiv:2310.10644;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2310.10644v1