MagicBrush : A Manually Annotated Dataset for Instruction-Guided Image Editing
Kai Zhang1 Lingbo Mo1∗ Wenhu Chen2 Huan Sun1 Yu Su1 1The Ohio State University 2University of Waterloo {zhang.13253, mo.169, https://osu-nlp-group.github.io/MagicBrush Equal Contribution.
Abstract
Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which contains a high volume of noise. Thus, they still require lots of manual tuning to produce desirable outcomes in practice. To address this issue, we introduce MagicBrush (https://osu-nlp-group.github.io/MagicBrush/), the first large-scale, manually annotated dataset for instruction-guided real image editing that covers diverse scenarios: single-turn, multi-turn, mask-provided, and mask-free editing. MagicBrush comprises over 10K manually annotated triplets (source image, instruction, target image), which supports training large-scale text-guided image editing models. We fine-tune InstructPix2Pix on MagicBrush and show that the new model can produce much better images according to human evaluation. We further conduct extensive experiments to evaluate current image editing baselines from multiple dimensions including quantitative, qualitative, and human evaluations. The results reveal the challenging nature of our dataset and the gap between current b
中文速览
用自然语言指令来编辑真实图片是个很实用的需求,但现有方法要么靠零样本推理需要大量手动调参,要么依赖自动合成的嘈杂数据集训练,效果都差强人意。为此,研究团队构建了 MagicBrush——首个大规模、人工标注的指令引导图像编辑数据集,包含超过 1 万组(源图像、编辑指令、目标图像)三元组,覆盖单轮、多轮、有掩码和无掩码等多种真实编辑场景。他们招募 Amazon Mechanical Turk 工人,让其借助 DALL-E 2 平台反复尝试直到产出满意的编辑结果,并经过严格筛选和质量把关,最终标注质量在一致性和图像品质两项人工评分上均达到5分中的4分左右。在 MagicBrush 上微调 InstructPix2Pix 之后,模型在人类偏好评测中显著优于各类基线方法,同时全面的定量与定性实验也表明现有方法与真实世界编辑需求之间仍存在明显差距,这一高质量数据集的发布有望推动该领域更先进模型的发展。
原文 arXiv:2306.10012;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2306.10012v3