|国家预印本平台
首页|A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging

A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging

A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging

来源:Arxiv_logoArxiv
英文摘要

Polyak-Ruppert averaging is a widely used technique to achieve the optimal asymptotic variance of stochastic approximation (SA) algorithms, yet its high-probability performance guarantees remain underexplored in general settings. In this paper, we present a general framework for establishing non-asymptotic concentration bounds for the error of averaged SA iterates. Our approach assumes access to individual concentration bounds for the unaveraged iterates and yields a sharp bound on the averaged iterates. We also construct an example, showing the tightness of our result up to constant multiplicative factors. As direct applications, we derive tight concentration bounds for contractive SA algorithms and for algorithms such as temporal difference learning and Q-learning with averaging, obtaining new bounds in settings where traditional analysis is challenging.

Sajad Khodadadian、Martin Zubeldia

计算技术、计算机技术自动化基础理论

Sajad Khodadadian,Martin Zubeldia.A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging[EB/OL].(2025-05-27)[2025-06-16].https://arxiv.org/abs/2505.21796.点此复制

评论