PoS (Progression of States (PoS)) runs an LLM agent against an explicit, structured “belief state” of the world plus unresolved task gaps, auditing each update for consistency and detecting when the agent is spinning its wheels (Belief Trapping) so it can inject a targeted recovery constraint before the budget runs out.
You’ve built a long-running agent. It uses tools, accumulates observations, and after 30 steps it’s confidently acting on a stale or self-contradictory picture of the world. Classic example from the paper: the agent’s notes say a microwave is both open and closed, so it can’t decide whether to open it before placing a mug. Or it keeps calling diagnostic tools that don’t actually narrow down which service failed.
The standard fix is better memory: compress the trajectory, summarize by subgoal, or retrieve relevant chunks. Methods like ReAct, ACON, HiAgent, and PACE all live in this “memory as context” family. The paper’s argument is that a reorganized transcript is still just evidence, not a current estimate of what’s true right now. The closest prior work, LongHorizon-Harness, does maintain an audited task state, but it depends on independently verifiable artifacts and struggles when progress means refining uncertain hypotheses (think clinical diagnosis or root cause analysis).
The core idea: at every step the agent carries an explicit object $B_t$ that bundles four things. A world state $W_t$ written as Entities, their States, and Relations between them (natural language, with provenance and confidence tags). A goal $G$. Two gap sets: epistemic gaps (things the agent still needs to learn) and achievement gaps (states of the world the agent still needs to produce). The agent picks one active gap to focus on, then samples its next action conditioned on the full belief plus that focus.
Two guardrails sit on top. First, a Belief Sentinel audits every proposed belief update for internal contradictions (microwave open-and-closed) and for contradictions against the latest observation, and asks the agent to revise before committing. Second, PoS computes a belief health score $\mathcal{H}_t$ over a sliding window of the last K validated transitions. Plain-English version: health drops when a gap has stayed unresolved for the whole window AND either no step is making recorded progress OR the agent keeps cycling back to the same world state. The score is 1 - p * max(S, R) where p is gap persistence, S is stagnation rate, R is recurrence rate. The paper deliberately avoids a geometric mean (which would miss stagnation-without-recurrence) and a plain max (which would flag any unfinished gap as trapped).
When health drops below a threshold, PoS does factorized diagnosis on two axes. The dynamics axis is Static (world unchanged), Cycle (periodic revisits), or Drift (world changes but not toward the active gap). The gap axis picks whichever gap type (epistemic vs achievement) has been more persistently blocked. The combination determines a recovery constraint added to the next action prompt: suppress the ineffective transition, break the detected cycle, re-anchor to the gap, require new discriminative evidence, or require a task-relevant state change.
B = init_belief(goal, obs0)
active_gap = pick_gap(B); C = set()
for t in range(T):
a = agent.act(B, active_gap, C)
obs = env.step(a)
B_tilde = agent.update(B, a, obs)
B = belief_sentinel.audit_and_revise(B, B_tilde, a, obs)
u = label_progress(B_prev, B, active_gap) # 0 or 1
active_gap = advance_if_resolved(B, active_gap)
if window_full(t):
H = health(window) # Eq. 5
if H <= theta_H:
rho = classify_dynamics(window) # Static/Cycle/Drift
X = pick_blocked_gap_type(P_E, P_A)
C = pattern_constraint(rho) | gap_constraint(X, B, active_gap)
Evaluated on four benchmarks with three backbones (Qwen3.7-Plus, Kimi K3, GLM-5.3). Two execution benchmarks: ALFWorld and LOCA-Bench. Two diagnosis benchmarks: RCA-100 (cloud incident root-cause analysis) and ClinDiag (clinical case diagnosis).
•
PoS wins all 12 cells (4 benchmarks × 3 backbones). The paper’s framing of the margins: relative gains over the strongest baseline reach 22.68% on ALFWorld, 37.89% on RCA-100 joint accuracy, and 11.31% on ClinDiag.
•
Memory-based baselines are inconsistent. PACE beats Raw Trajectory on ALFWorld but trails it by 24–28 points on LOCA-Bench. On ClinDiag with GLM-5.3, no context-management baseline beats just dumping the raw trajectory in. The authors read this as: modern long-context LLMs don’t always benefit from compression, because information loss can outweigh redundancy reduction.
•
Ablations show both guardrails matter, differently. Dropping Consistency Validation hits execution hardest (down 14.93 points on ALFWorld, 11.81 on LOCA-Bench) but barely touches ClinDiag (0.33–0.66 points). Dropping trapping diagnosis hurts RCA-100 by 4.85–6.79 points and ClinDiag by 2.65–3.97. So validation is critical where state changes compound, and recovery is critical where investigations stall.
•
Trapping patterns differ by domain. Cycles dominate in ALFWorld (agent revisits rooms), drift in LOCA-Bench and RCA-100, static stagnation in ClinDiag (78.25% of trapped episodes). A variant that replaces factorized recovery with a generic “please recover” prompt loses on all four benchmarks.
•
Context scaling (LOCA-Bench, 8K→256K). PoS stays roughly flat from 96K to 256K and beats the strongest baseline by 10.67–16.00 points at 256K.
•
Cost. On RCA-100 with Qwen3.7-Plus, PoS cuts Task Agent tokens by 20.9% but total token spend rises to 5.06× Raw Trajectory, because the Belief and Sentinel components each burn hundreds of thousands of tokens per episode.
•
If you run long-horizon agents and your current failure mode is “it loops” or “it acts on stale facts,” the specific contribution worth borrowing is the health signal: track unresolved-gap persistence AND (stagnation OR recurrence) together over a sliding window. The paper shows generic “are you stuck?” prompts underperform diagnosis-conditioned constraints, so if you add trap detection, also classify the trap type before writing the recovery prompt.
•
The Entity-State-Relation belief format is domain-agnostic by construction, but you’re paying for it. The 5× token cost is real, and the authors flag it as the top limitation. If budget matters more than peak accuracy, the Consistency Validation ablation gives a cheaper midpoint on diagnosis tasks (saves 35.6% of tokens, loses 7.76 points on RCA-100 joint accuracy). On execution tasks that tradeoff looks worse.
•
Don’t read these results as proof that explicit belief helps every agent. The paper’s own category breakdown shows gains concentrate where execution requires tracking state changes across steps (ALFWorld Transform) or where diagnosis requires combining signals to separate hypotheses (RCA-100 fault triage). On ALFWorld Place, PoS matches but doesn’t beat the strongest baseline.
•
The authors are explicit that PoS doesn’t fix missing domain knowledge or weak tool-use skill. Many ClinDiag failures happen after the agent collects the right evidence. If your bottleneck is “the model doesn’t know medicine,” belief maintenance won’t rescue you.
•
Code is at GitHub; project page at the authors’ site.
•
The reported gains assume you can afford 3-5× the token spend of a Raw Trajectory baseline. The Belief and Sentinel components each run as separate LLM calls per step.
•
All backbones are strong recent models accessed through one provider’s API. The paper does not test weaker open models, where generating structured beliefs and auditing them reliably may be harder.
•
The medical-diagnosis judge on ClinDiag is itself Qwen3.7-Plus, held fixed across backbones. That makes within-benchmark comparisons clean but means absolute ClinDiag scores inherit that judge’s biases.
•
Hyperparameters (window size K=8, health threshold 0.25, recurrence threshold 0.75, etc.) are fixed per benchmark and not swept per backbone. Transferring PoS to a new domain likely requires retuning these.
•
Comparisons to several belief-oriented methods (Tru-POMDP, Graph of States) are omitted because those methods are tied to specific task structures. The strongest state-tracking baseline tested is LongHorizon-Harness.