A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
Deep theoretical contribution advancing the mathematical understanding of distributional reinforcement learning algorithms, essential for researchers in RL theory.
AI Summary
This paper provides a rigorous finite-sample analysis for quantile temporal-difference learning, establishing convergence rates and separating local stochastic fluctuation from global sample complexity in distributional RL.
Excerpt
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean
