Stochastic subgradient method converges on tame functions
Damek Davis Thanks: School of Operations Research and Information Engineering, Cornell University, Ithaca, NY 14850, USA; people.orie.cornell.edu/dsd95/. Dmitriy Drusvyatskiy Thanks: Department of Mathematics, University of Washington, Seattle, WA 98195; www.math.washington.edu/$∼$ddrusv. Research of Drusvyatskiy was supported by the AFOSR YIP award FA9550-15-1-0237 and by the NSF DMS 1651851 and CCF 1740551 awards. Sham Kakade Thanks: Departments of Statistics and Computer Science, University of Washington, Seattle, WA 98195; homes.cs.washington.edu/$∼$sham/. Sham Kakade acknowledges funding from the Washington Research Foundation Fund for Innovation in Data-Intensive Discovery and the NSF CCF 1740551 award. Jason D. Lee Thanks: Data Science and Operations Department, Marshall School of Business, University of Southern California, Los Angeles, CA 90089; www-bcf.usc.edu/$∼$lee715. JDL acknowledges funding from the ARO MURI Award W911NF-11-1-0303.
Abstract
This work considers the question: what convergence guarantees does the stochastic subgradient method have in the absence of smoothness and convexity? We prove that the stochastic subgradient method, on any semialgebraic locally Lipschitz function, produces limit points that are all first-order stationary. More generally, our result applies to any function with a Whitney stratifiable graph. In particular, this work endows the stochastic subgradient method, and its proximal extension, with rigorous convergence guarantees for a wide class of problems arising in data science—including all popular deep learning architectures.
原文 arXiv:1804.07795;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1804.07795v3