Get Started
Research questionHow can reinforcement-learning planning optimize nonlinear objectives over visitation measures?Planning-as-inference is often formulated around linear rewards, leaving the treatment of nonlinear functionals of state-action visitation unclear. The interpretation of optimization updates and temporal-difference errors also needs to extend beyond that linear setting.
Machine Learning
Reinforcement Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.The Dually Flat Geometry of Planning as InferenceThe source studies reinforcement-learning planning through a resetting planning process whose stationary distribution is a visitation measure. It develops a dually flat information-geometric characterization, natural-gradient updates for nonlinear visitation objectives, and an interpretation of temporal-difference error as a marginal-utility estimate; the evidence is theoretical and also discusses implications for theoretical neuroscience.research paper · Sep 3, 2026
Related questions
How can reinforcement-learning data acquisition be optimized when interactions are costly, slow, and human-mediated?How can functional bilevel optimization adapt online as learning objectives change over time?How can distributed VLA reinforcement learning coordinate variable-latency simulation, inference, and optimization?How can reinforcement learning post-training for diffusion models avoid objective mismatch and high-variance updates?
Home
Topics
Search
Library