Semantic Score Distillation Sampling for Compositional Text-to-3D Generation
Ling Yang Thanks: Contributed equally. Zixiang Zhang Affiliation: Peking University Junlin Han Affiliation: University of OxfordProject: https://github.com/YangLing0818/SemanticSDS-3D Bohan Zeng Affiliation: Peking University Runjia Li Affiliation: University of OxfordProject: https://github.com/YangLing0818/SemanticSDS-3D Philip Torr Affiliation: University of OxfordProject: https://github.com/YangLing0818/SemanticSDS-3D Wentao Zhang Thanks: Corresponding authors: Affiliation: Peking University
Abstract
Generating high-quality 3D assets from textual descriptions remains a pivotal challenge in computer graphics and vision research. Due to the scarcity of 3D data, state-of-the-art approaches utilize pre-trained 2D diffusion priors, optimized through Score Distillation Sampling (SDS). Despite progress, crafting complex 3D scenes featuring multiple objects or intricate interactions is still difficult. To tackle this, recent methods have incorporated box or layout guidance. However, these layout-guided compositional methods often struggle to provide fine-grained control, as they are generally coarse and lack expressiveness. To overcome these challenges, we introduce a novel SDS approach, Semantic Score Distillation Sampling (SemanticSDS), designed to effectively improve the expressiveness and accuracy of compositional text-to-3D generation. Our approach integrates new semantic embeddings that maintain consistency across different rendering views and clearly differentiate between various objects and parts. These embeddings are transformed into a semantic map, which directs a region-specific SDS process, enabling precise optimization and compositional generation. By leveraging explicit s
原文 arXiv:2410.09009;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2410.09009v1