Research questionHow can offline reinforcement learning improve policies beyond dataset support while keeping value estimates reliable under distribution shift?Offline RL learns from a fixed dataset, so policy changes toward unsupported actions can expose the critic to distribution shift and unreliable value estimates. Staying too close to observed behavior, however, can limit meaningful policy improvement.