Feature Space Augmentation for Long-Tailed Data
Peng Chu Thanks: Work was done at GE Research. Affiliation: Temple University, USA E-mail Xiao Bian* Affiliation: Google Inc., USA E-mail Shaopeng Liu Affiliation: GE Research, USA E-mail Haibin Ling Affiliation: Temple University, USA E-mail Affiliation: Stony Brook University, USA E-mail
Abstract
Real-world data often follow a long-tailed distribution as the frequency of each class is typically different. For example, a dataset can have a large number of under-represented classes and a few classes with more than sufficient data. However, a model to represent the dataset is usually expected to have reasonably homogeneous performances across classes. Introducing class-balanced loss and advanced methods on data re-sampling and augmentation are among the best practices to alleviate the data imbalance problem. However, the other part of the problem about the under-represented classes will have to rely on additional knowledge to recover the missing information.
原文 arXiv:2008.03673;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2008.03673v1