Temporal Pyramid Pooling Based Convolutional Neural Network for Action Recognition
Peng Wang Yuanzhouhan Cao Thanks: P. Wang’s contribution was made when visiting the University of Adelaide. The first two authors made equal contributions to this work. Chunhua Shen Lingqiao Liu Heng Tao Shen Thanks: P. Wang and H. T. Shen are with School of Information Technology and Electrical Engineering, The University of Queensland, Australia (email: Thanks: Y. Cao, C. Shen, and L. Liu are with School of Computer Science, The University of Adelaide, Australia (email: {yuanzhouhan.cao, chunhua.shen, Thanks: C. Shen is also with Australian Centre for Robotic Vision, Australia.
Abstract
Encouraged by the success of Convolutional Neural Networks (CNNs) in image classification, recently much effort is spent on applying CNNs to video based action recognition problems. One challenge is that video contains a varying number of frames which is incompatible to the standard input format of CNNs. Existing methods handle this issue either by directly sampling a fixed number of frames or bypassing this issue by introducing a 3D convolutional layer which conducts convolution in spatial-temporal domain.
原文 arXiv:1503.01224;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1503.01224v2