TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering
Yushi Hu Affiliation: University of Washington Benlin Liu Affiliation: University of Washington Jungo Kasai Affiliation: University of Washington Yizhong Wang Affiliation: University of Washington Mari Ostendorf Affiliation: University of Washington Ranjay Krishna Affiliation: University of Washington Affiliation: Allen Institute for AIhttps://tifa-benchmark.github.io/ Noah A. Smith Affiliation: University of Washington Affiliation: Allen Institute for AIhttps://tifa-benchmark.github.io/
Abstract
Despite thousands of researchers, engineers, and artists actively working on improving text-to-image generation models, systems often fail to produce images that accurately align with the text inputs. We introduce TIFA (Text-to-Image Faithfulness evaluation with question Answering), an automatic evaluation metric that measures the faithfulness of a generated image to its text input via visual question answering (VQA). Specifically, given a text input, we automatically generate several question-answer pairs using a language model. We calculate image faithfulness by checking whether existing VQA models can answer these questions using the generated image. TIFA is a reference-free metric that allows for fine-grained and interpretable evaluations of generated images. TIFA also has better correlations with human judgments than existing metrics. Based on this approach, we introduce TIFA v1.0, a benchmark consisting of 4K diverse text inputs and 25K questions across 12 categories (object, counting, etc.). We present a comprehensive evaluation of existing text-to-image models using TIFA v1.0 and highlight the limitations and challenges of current models. For instance, we find that current
原文 arXiv:2303.11897;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2303.11897v3