Get Started
Home
Topics
Search
Library
Topic · 31 recaps
Reinforcement Learning
Learning from reward signals through trial and interaction with an environment. Spans classic RL, RLHF, and modern post-training methods for language and agent models.
...
Sort
Newest
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
Evaluation · Aug 27
0
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
Agents · Aug 25
0
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
Multimodal · Aug 20
0
EnvHarness: Awakening Static Worlds for Agent Learning
Agents · Aug 20
0
DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents
Agents · Aug 19
0
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Reasoning · Aug 17
0
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Agents · Aug 10
0
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Agents · Aug 6
0
Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
Multimodal · Aug 6
0
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
LLM Training · Aug 6
0