Research questionHow can we learn an effective LQR controller from unknown dynamics without a stable initial policy?Unknown system dynamics make it difficult to learn a controller with limited interaction data. Requiring a stabilizing initial policy can also exclude systems where no such controller is available beforehand.