Characterization of Excess Risk for Locally Strongly Convex Population Risk

Yi, Mingyang; Wang, Ruoyu; Ma, Zhi-Ming

Computer Science > Machine Learning

arXiv:2012.02456 (cs)

[Submitted on 4 Dec 2020 (v1), last revised 8 Oct 2022 (this version, v4)]

Title:Characterization of Excess Risk for Locally Strongly Convex Population Risk

Authors:Mingyang Yi, Ruoyu Wang, Zhi-Ming Ma

View PDF

Abstract:We establish upper bounds for the expected excess risk of models trained by proper iterative algorithms which approximate the local minima. Unlike the results built upon the strong globally strongly convexity or global growth conditions e.g., PL-inequality, we only require the population risk to be \emph{locally} strongly convex around its local minima. Concretely, our bound under convex problems is of order $\tilde{\cO}(1/n)$. For non-convex problems with $d$ model parameters such that $d/n$ is smaller than a threshold independent of $n$, the order of $\tilde{\cO}(1/n)$ can be maintained if the empirical risk has no spurious local minima with high probability. Moreover, the bound for non-convex problem becomes $\tilde{\cO}(1/\sqrt{n})$ without such assumption. Our results are derived via algorithmic stability and characterization of the empirical risk's landscape. Compared with the existing algorithmic stability based results, our bounds are dimensional insensitive and without restrictions on the algorithm's implementation, learning rate, and the number of iterations. Our bounds underscore that with locally strongly convex population risk, the models trained by any proper iterative algorithm can generalize well, even for non-convex problems, and $d$ is large.

Comments:	The first two authors contribute equally to this paper
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2012.02456 [cs.LG]
	(or arXiv:2012.02456v4 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2012.02456

Submission history

From: Mingyang Yi [view email]
[v1] Fri, 4 Dec 2020 08:24:50 UTC (3,573 KB)
[v2] Fri, 29 Jan 2021 11:35:16 UTC (3,540 KB)
[v3] Wed, 26 May 2021 08:14:43 UTC (4,004 KB)
[v4] Sat, 8 Oct 2022 03:09:29 UTC (1,462 KB)

Computer Science > Machine Learning

Title:Characterization of Excess Risk for Locally Strongly Convex Population Risk

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Characterization of Excess Risk for Locally Strongly Convex Population Risk

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators