To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression
Yitian Yuan Affiliation: Tsinghua-Berkeley Shenzhen Institute, Tsinghua University, China Tao Mei Affiliation: JD AI Research, China Wenwu Zhu Affiliation: Tsinghua-Berkeley Shenzhen Institute, Tsinghua University, China Affiliation: Department of Computer Science and Technology, Tsinghua University,
Abstract
We have witnessed the tremendous growth of videos over the Internet, where most of these videos are typically paired with abundant sentence descriptions, such as video titles, captions and comments. Therefore, it has been increasingly crucial to associate specific video segments with the corresponding informative text descriptions, for a deeper understanding of video content. This motivates us to explore an overlooked problem in the research community — temporal sentence localization in video, which aims to automatically determine the start and end points of a given sentence within a paired video. For solving this problem, we face three critical challenges: (1) preserving the intrinsic temporal structure and global context of video to locate accurate positions over the entire video sequence; (2) fully exploring the sentence semantics to give clear guidance for localization; (3) ensuring the efficiency of the localization method to adapt to long videos. To address these issues, we propose a novel Attention Based Location Regression (ABLR) approach to localize sentence descriptions in videos in an efficient end-to-end manner. Specifically, to preserve the context information, ABLR fi
原文 arXiv:1804.07014;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1804.07014v4