Research questionHow can offline goal-conditioned reinforcement learning learn reliable values for long-horizon tasks without compounding overestimation?In offline goal-conditioned reinforcement learning, long-range value estimates depend on shorter-range estimates that may already be inaccurate. Repeated max-based backups can amplify those errors across long-horizon tasks.