Get Started
Home
Topics
Search
Library
7 min read · Agents · Evaluation · Added Oct 7 · Paper published Sep 29, 2026

Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

Source: research paper via Hugging Face Daily Papers
Proactive LLM agents fail not by doing wrong work but by doing right work at wrong moments or wrong intervention depths. Proactivity-Gym scores task, timing, and trust separately: top model hits 65/100 on task but only 52% depth-match, and LLM judges rate trust where humans withdraw it.
TL;DR
This paper argues that proactive LLM agents must be designed and evaluated along three joint axes (task capability, temporal allocation, and trust), and introduces a multi-day simulation testbed where even the strongest tested agent clears only ~65/100 on task quality and ~52% on matching the user’s preferred intervention depth.
Why It Matters
Imagine you’ve wired an LLM assistant into your calendar, email, and shell. Between your explicit requests it has idle compute. The tempting move is to let it work ahead: draft replies, pre-run experiments, book that review slot. The failure mode is that correct proactive work can still be unwanted. It interrupts you mid-focus, uses GPU you needed, or quietly executes something you wanted to approve first. Each misstep costs review time and, more importantly, trust.
Most existing proactive-agent benchmarks score whether the agent predicted the right need or produced the right artifact. That collapses several distinct failures into one number. An agent can produce a great summary at the wrong moment, or produce a mediocre summary at a perfect moment, and the paper’s claim is that these should not be scored the same way. The authors position this against prior proactive-agent work (notably Mixed-initiative interaction systems and Sleep-time compute) that each addressed one piece but not the joint problem.
How It Works
The contribution is a framework, not an algorithm. It has three layers.
First, the 3T principles. Task Capability (TC) is whether the agent correctly anticipates a need and does the work well. Temporal Allocation (TA) is whether the agent uses compute at the right time, distinguishing interaction time (user is active, compute is contested) from sleep time (user is away, compute is free). Trust (TR) is whether the user remains willing to rely on the agent, measured partly through intervention depth: does the agent prepare, suggest, or execute without asking.
Second, a design space with five dimensions: task scope (within or outside the current task), anticipation horizon (now, later today, beyond), activation trigger (user, event, or agent-initiated), processing timing (interaction or sleep time), and intervention depth. Each 3T objective informs specific dimensions.
Third, Proactivity-Gym, a simulation testbed with 10 hand-crafted scenarios spanning 7-10 simulated days each. A simulated clock advances; tool state persists; events fire on schedule; a persona-conditioned user simulator (three personas per scenario: Reviewer, Planner, Operator) grants or denies approval based on a deterministic state machine. The agent’s control loop at each step looks roughly like:
while sim_time < scenario_end: obs = env.observe() # tool state, notifications, user msgs situation = build_situation(context, obs) # user goals, trust, deadlines action = agent.choose(situation, context) # action \u2208 {prepare, suggest, execute, defer, noop} if action.depth == "execute" and needs_approval(action): reply = user_sim.respond(action) # may approve/deny/defer context = update(context, action, env.step(action))
Scoring: TC combines rule-based checks (evidence retrieved, outputs produced, constraints met) with LLM-as-a-Judge scores for semantic and completion quality. TA is binary per scenario, requiring the agent to both prioritize the urgent task and explicitly defer the competing one. TR has two sub-metrics: TR-D (fraction of tasks where chosen intervention depth matches the persona’s expected depth) and TR-J (LLM judges rate five trust constructs: understandability, technical competence, reliability, personal attachment, faith).
What They Found
Across 23 model-harness combinations (nine models including Claude Opus 5, GPT-5.6-sol, Qwen 3.5 variants, Gemma 4 variants; three harnesses: OpenClaw, Claude Code, Codex), proactive performance was uneven across 3T.
•
Big gaps even at the top. Claude Opus 5 led with TC = 65.1/100, TA = 51.7%, TR-D = 52.4%. Every other model landed below 20% on TA with TR-D between 39-51%. Larger models in a family did better across all four metrics.
•
Deferring is the hard part. On TA, models picked the right urgent task ~44% of the time but explicitly deferred the competing task only ~11% of the time. Silence doesn’t count as deferral, which the authors argue reflects a real capability gap.
•
TC does not drag TA or TR-D with it. Run-level correlation of TC with TA is r = 0.31 and with TR-D is r = 0.34. TC correlates more with TR-J (r = 0.70), which the authors read as evidence that LLM judges confuse “did good work” with “earned trust.”
•
More reasoning helps TC and TA, not trust. For GPT 5.6-Sol, raising reasoning effort lifted TC from 50.0 to 58.9 but left both trust metrics flat.
•
Harness choice matters, non-uniformly. Swapping Claude Code for OpenClaw on open-weight models shifted TC by -11.1 to +0.5 depending on the model.
•
Human study (30 participants) diverges from LLM judges. Participants preferred sleep-time assistance that needed correction over immediate correct assistance that competed with ongoing work (97.8% vs 26.7% acceptance). After a single misaligned intervention (correct outcome, wrong intervention depth), mean trust dropped 1.86 points; recovery after returning to aligned behavior was only 1.27 points. In an aligned\u2192misaligned\u2192aligned sequence, trust ended at 3.43 versus a starting 4.43, even though the final action itself was rated 4.71.
The gap between TR-J (LLM-rated trust) and human trust ratings is the paper’s sharpest finding: LLM judges give Claude Opus 5 high trust scores (4.72 understandability, 4.60 competence) despite its 52.4% intervention-depth match, while humans penalize that mismatch heavily.
What’s Useful
•
If you’re building a proactive agent and only tracking “did it do the right task,” you’re probably missing where it actually fails users. The paper’s practical advice is to separately instrument when the agent acts and how deep it intervenes. The TR-D metric (match between chosen depth and user preference) is cheap to compute if you have per-user policies and worth logging in production.
•
If you’re using an LLM-as-a-Judge setup to score agent behavior, treat trust-like judgments with suspicion. The paper’s evidence (TC correlates r=0.70 with judge-rated trust but humans react very differently to intervention misalignment) suggests judges will approve agents that users would silently stop trusting. Worth testing: run an A/B with real users on the subset of trajectories where TR-D and TR-J disagree.
•
Loss aversion for trust is worth designing around. The human study shows one misaligned action costs more than one aligned action gains, and recovery is incomplete. If you’re deciding between “ask for approval and risk being annoying” vs “execute and risk being wrong about depth,” the paper’s evidence leans toward asking, especially early in a user relationship.
•
The scheduling angle is underused: participants accepted imperfect sleep-time work far more than correct work that competed for their attention. If your agent has any notion of user availability (calendar, activity, OS idle), using that to defer non-urgent work is likely higher-leverage than making the work itself better.
•
Code and gym data are promised “upon publication,” not yet linked in the paper. Project page.
Caveats
The 10 scenarios are hand-authored and the user simulator is an LLM following a deterministic approval state machine, so results reflect that simulator’s behavior, not real users. TA is reduced to a binary now/later decision on scripted conflicts; real scheduling involves variable task durations, changing compute budgets, and preemption, which the authors flag as future work. TR-J uses LLM judges with known calibration differences (Qwen vs Gemini averaged); the human study (n=30, mostly students and professionals familiar with LLMs) is scenario-based rather than longitudinal real-world use. The reference list contains several citations dated 2026 that the reader may not be able to verify.
Topics
Agents
Evaluation
Agents
Evaluation
Up next in Agents
World Editing: Intervening on Executable Worlds at Increasing Depth
Optimizing the Optimizer: Language Models Discover Faster Molecular Relaxation
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents231 episodes
Evaluation167 episodes