Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models
Shihao ZhaoThe University of Hong HaoThe University of Hong Thanks: Corresponding Author, $†$ Intern at Microsoft Kwan-Yee K. WongThe University of Hong
Abstract
Text-to-Image diffusion models have made tremendous progress over the past two years, enabling the generation of highly realistic images based on open-domain text descriptions. However, despite their success, text descriptions often struggle to adequately convey detailed controls, even when composed of long and complex texts. Moreover, recent studies have also shown that these models face challenges in understanding such complex texts and generating the corresponding images. Therefore, there is a growing need to enable more control modes beyond text description. In this paper, we introduce Uni-ControlNet, a unified framework that allows for the simultaneous utilization of different local controls (e.g., edge maps, depth map, segmentation masks) and global controls (e.g., CLIP image embeddings) in a flexible and composable manner within one single model. Unlike existing methods, Uni-ControlNet only requires the fine-tuning of two additional adapters upon frozen pre-trained text-to-image diffusion models, eliminating the huge cost of training from scratch. Moreover, thanks to some dedicated adapter designs, Uni-ControlNet only necessitates a constant number (i.e., 2) of adapters, reg
原文 arXiv:2305.16322;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2305.16322v3