Non-parametric Stochastic Approximation with Large Step-sizes
Aymeric [ Francis [ Département d’Informatique de l’Ecole Normale Supérieure, Paris, France SIERRA Project-Team 23, avenue d’Italie 75013 Paris, France E-mail: e2
Abstract
We consider the random-design least-squares regression problem within the reproducing kernel Hilbert space (RKHS) framework. Given a stream of independent and identically distributed input/output data, we aim to learn a regression function within an RKHS $\mathcal{H}$ , even if the optimal predictor (i.e., the conditional expectation) is not in $\mathcal{H}$ . In a stochastic approximation framework where the estimator is updated after each observation, we show that the averaged unregularized least-mean-square algorithm (a form of stochastic gradient descent), given a sufficient large step-size, attains optimal rates of convergence for a variety of regimes for the smoothnesses of the optimal prediction function and the functions in $\mathcal{H}$ .
中文速览
如何在再生核希尔伯特空间(RKHS)中用在线随机梯度下降做非参数回归、同时又不依赖额外正则化超参数,并实现理论上最优的收敛速率,是该领域长期悬而未决的问题。本文证明:只要把步长设置得足够大(而非传统的递减步长),再对迭代结果做均值平均(即averaged least-mean-square算法),无需额外正则化项,就能在多种光滑度假设下同时达到最优收敛速率——即便真实条件期望并不属于所用的RKHS。理论结果覆盖有限样本量和在线两种场景,给出了步长的显式选取公式,并在合成样条平滑实验中得到了验证。这一工作从根本上简化了核方法的在线学习流程:用单一超参数(步长)替代了传统方法中步长与正则化参数的双重调节,同时解决了此前文献中明确提出的开放问题。
原文 arXiv:1408.0361;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1408.0361v3