Towards Robust Evaluations of Continual Learning
Sebastian Farquhar Yarin Gal
Abstract
Experiments used in current continual learning research do not faithfully assess fundamental challenges of learning continually. Instead of assessing performance on challenging and representative experiment designs, recent research has focused on increased dataset difficulty, while still using flawed experiment set-ups. We examine standard evaluations and show why these evaluations make some continual learning approaches look better than they are. We introduce desiderata for continual learning evaluations and explain why their absence creates misleading comparisons. Based on our desiderata we then propose new experiment designs which we demonstrate with various continual learning approaches and datasets. Our analysis calls for a reprioritization of research effort by the community.
中文速览
持续学习(continual learning)领域长期存在一个被忽视的问题:研究者普遍采用设计缺陷的实验来评估算法,导致某些方法看起来表现良好,实则掩盖了严重的短板。作者系统梳理了当前主流评估方案的五大缺陷,并据此提出五条"评估基准准则"(desiderata),涵盖任务间输入相似性、输出向量设置、任务标识可用性、禁止回访旧数据以及任务数量规模等关键维度。通过在多个数据集上对EWC、SI、VCL等主流"先验聚焦"(prior-focused)方法进行重新测试,作者发现:一旦同时满足这五条准则,目前没有任何先验聚焦方法能够真正成功;换言之,现有评估体系系统性地高估了这类算法的实际能力。这项工作呼吁整个领域在追求更复杂数据集之前,先把实验设计本身做对,否则再难的数据集也无法真实衡量持续学习的进步。
原文 arXiv:1805.09733;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1805.09733v3