Research questionHow can model-free robust Q-learning learn from one trajectory with second-moment targets and a noncontractive Bellman operator?A single transition does not provide an unbiased estimate of the conditional second moment needed by chi-square robust Bellman targets. Convergence is further complicated because the projected robust Bellman operator need not be contractive.