BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models
Hanjun Luo , Haoyu Huang , Ziye Deng , Ruizhe Chen , Xinfeng Li , Hewei Wang , Yingbin Jin , Yang Liu , Wenyuan Xu , Zuozhu Liu Zhejiang University, Technological University, Mellon University, University, Technological University, University, University, Corresponding author
Abstract
Text-to-Image (T2I) generative models are becoming increasingly crucial due to their ability to generate high-quality images, but also raise concerns about social biases, particularly in human image generation. Sociological research has established systematic classifications of bias. Yet, existing studies on bias in T2I models largely conflate different types of bias, impeding methodological progress. In this paper, we introduce BIGbench, a unified benchmark for Biases of Image Generation, featuring a carefully designed dataset. Unlike existing benchmarks, BIGbench classifies and evaluates biases across four dimensions to enable a more granular evaluation and deeper analysis. Furthermore, BIGbench applies advanced multi-modal large language models to achieve fully automated and highly accurate evaluations. We apply BIGbench to evaluate eight representative T2I models and three debiasing methods. Our human evaluation results by trained evaluators from different races underscore BIGbench’s effectiveness in aligning images and identifying various biases. Moreover, our study also reveals new research directions about biases with insightful analysis of our results. Our work is openly ac
中文速览
文生图(Text-to-Image)模型在生成人物图像时会放大社会偏见,但现有评测基准要么覆盖面窄、只关注职业偏见,要么把不同类型的偏见混为一谈,难以指导后续改进。为此,研究团队提出了 BIGbench,一个从"后天属性、受保护属性、偏见表现形式、偏见可见性"四个维度系统分类偏见的统一评测框架,并构建了包含 47,040 条提示语的数据集,涵盖职业、社会关系和个人特征等场景。评测流程采用经过微调的多模态大语言模型实现全自动、高精度打分,可同时衡量隐式生成偏见、显式生成偏见、忽视(ignorance)和歧视(discrimination)四类结果。研究团队用 BIGbench 对 8 个主流文生图模型和 3 种去偏方法进行了横向比较,并通过来自不同种族的人工评估者对 1,000 张图像进行验证,证实了该框架的有效性;其发现为构建更公平的 AI 图像生成系统提供了具体指引,相关资源已全部开源。
原文 arXiv:2407.15240;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2407.15240v6