Get Started
Research questionHow can model-free robust Q-learning learn from one trajectory with second-moment targets and a noncontractive Bellman operator?A single transition does not provide an unbiased estimate of the conditional second moment needed by chi-square robust Bellman targets. Convergence is further complicated because the projected robust Bellman operator need not be contractive.
Machine Learning
Reinforcement Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Finite-Time Convergence of Single-Trajectory Chi-Square Robust Q-Learning With Linear Function ApproximationThe evidence concerns an unknown nominal MDP, single-trajectory data, model-free robust Q-learning, chi-square uncertainty sets, and linear function approximation. It establishes a finite-time error bound relative to the optimal robust Q-function for discount factors in (0,1); a neural-network experiment illustrates the target in a continuous-state nonlinear-control task.research paper · Sep 3, 2026
Related questions
How can we learn an effective LQR controller from unknown dynamics without a stable initial policy?How can learned chaotic systems preserve long-term statistics and stability beyond short-term trajectory matching?How can learned models approximate trajectory similarity across multiple granularities without losing relative similarity rankings?How can closed-loop vision-language navigation learn effectively despite distribution shift and sparse micro-action rewards?
Home
Topics
Search
Library