A Pooling Approach to Modelling Spatial Relations for Image Retrieval and Annotation
Mateusz Malinowski Max Planck Institute for Informatics Saarbrücken, Germany Mario Fritz Max Planck Institute for Informatics Saarbrücken, Germany
Abstract
Over the last two decades we have witnessed strong progress on modeling visual object classes, scenes and attributes that have significantly contributed to automated image understanding. On the other hand, surprisingly little progress has been made on incorporating a spatial representation and reasoning in the inference process. In this work, we propose a pooling interpretation of spatial relations and show how it improves image retrieval and annotations tasks involving spatial language. Due to the complexity of the spatial language, we argue for a learning-based approach that acquires a representation of spatial relations by learning parameters of the pooling operator. We show improvements on previous work on two datasets and two different tasks as well as provide additional insights on a new dataset with an explicit focus on spatial relations.
中文速览
用空间关系来理解图像中的"左边""上方"之类的位置描述,一直是计算机视觉中少有人系统解决的难题。作者提出把心理学中的"空间模板"(spatial template)重新解释为一种以参考物体为中心的可学习空间池化(spatial pooling)操作,让模型通过数据自动学习每种空间介词(如"above""right of")对应的位置权重分布,而无需手工设计规则。在图像检索和图文对齐两类任务上,该方法在两个已有数据集上均达到或超过此前最好的结果,并在一个专门聚焦空间关系的新数据集上提供了更细致的分析。这项工作的意义在于:它为空间语言理解提供了一个统一、灵活且可端到端学习的框架,弥补了现有视觉理解系统在空间推理能力上的明显短板。
原文 arXiv:1411.5190;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1411.5190v2