Time-Uniform Self-Normalized Concentration for Discounted Least Squares: Limits and Corrections

Chronological Source Flow
Back

AI Fusion Summary

A new research paper on arXiv:2608.19643v1 examines self-normalized concentration inequalities used in bandit and reinforcement-learning analyses. The study challenges a weighted extension claiming time-uniform guarantees for discounted least-squares estimators in non-stationary problems. Using a scalar Gaussian counterexample, the authors demonstrate that the claimed bounded radius is crossed with probability one. Furthermore, they prove that for specific discount and regularization parameters, any deterministic anytime boundary must be at least of order R√sqrt(log(T/δ)) by horizon T.
Community Comments
Loading updates...
0