SafeActBench measures whether tool-using agents actually establish the evidence their consequential actions require, finding that strong static action-judgment coexists with much weaker interactive execution because agents stop investigating too early or act before prerequisites are confirmed.
You’ve built an agent that can call tools: issue a refund, update a database row, dispatch a courier. The usual way to tell if it worked is to check the endpoint. Did the refund go through? Did the right record change? That’s what most interactive benchmarks like \u03c4-bench or AppWorld score: milestones, final state, policy constraints.
The authors argue that endpoint correctness hides a serious failure mode. An agent can refund the correct charge for the correct amount and get a success receipt, without ever having verified in its trajectory that the charge was a duplicate or that it was refund-eligible. The task state happened to permit the action. The agent got lucky. If a later step depends on evidence the agent never actually gathered, that luck runs out.
So the question this paper asks is narrower than “did the task succeed”: did the agent establish, in observable tool interactions, the specific facts about the specific entity that its action required, before it acted? And when actions chain, did later actions bind to results the earlier actions actually produced?
The core idea is to score the evidence-to-action chain, not the endpoint. For every consequential action, the benchmark specifies required evidence (target entity, current state, amount, authorization, etc.), the exact action, and any downstream dependencies. A deterministic evaluator replays the trajectory and checks: was each required fact established, from a tool observation bound to the right entity, before the action fired? Reading charge C1’s amount does not establish C2’s amount even if the numbers match. This is called provenance-bound evidence tracking via an Evidence Ledger.
The benchmark, SafeActBench, has 656 cases across six domains (customer ops, infra, legal/finance, research, smart home, healthcare) and five protocols of increasing demand:
•
Legacy: static Allow/Block/Defer judgment on a fixed candidate action.
•
V0: investigated non-action. The agent must gather enough to justify why not to act, then stop.
•
V1: one correct consequential action, with all prerequisites established first.
•
V2: linear multi-action workflow where later actions consume earlier results.
•
V3: dependency graph of actions, any topological order allowed.
Scoring is binary per episode (exact_case_success) and uses no LLM judge. The evaluator checks tools, targets, arguments, result propagation, dependency order, and terminal state.
for action in trajectory.consequential_actions:
for req in spec.requirements(action):
if not ledger.established(req, entity=action.target,
before=action.timestamp):
fail("unsupported action")
if action.depends_on:
if not uses_actual_result(action, action.depends_on):
fail("dependency not bound to real predecessor result")
check_terminal_state(trajectory)
Ten configurations were tested: five model families (Claude Opus 5, GPT-5.6, DeepSeek-V4, Qwen3.8, GLM-5.2) each paired with their family harness (Claude Code, Codex, DSH, Qwen Code, ZCode) and with a shared ReAct harness in Inspect AI.
•
Static judgment vastly overstates agentic ability. GLM–ZCode scores 97.7% on Legacy but 12.1% on V2 and 34.1% on V3. DeepSeek–DSH hits 96.5% Legacy and sits around 60% on V1–V3. On a matched 100-case cohort, static accuracy was ≥95% while interactive ECS was ≤52%, so this isn’t a case-composition artifact.
•
Failures happen before execution, not during it. Agents often stop with incomplete investigation (BSR 21.7–62.9% on V0) or act before evidence is complete (PAR 37.0–66.9% on V1 among episodes that acted). Conditional on acting after investigation completed, action success is 93.2–100% for nine of ten configs. The execution step itself is largely reliable; the decision of when to execute is where things break.
•
Harness changes the shape of failures, not just the rate. GLM under Inspect vs. ZCode trades early-action failures for no-action failures with almost identical total failure rate. Harness-paired deltas range from +4.4 pp (DeepSeek with DSH) to −6.8 pp (GLM with ZCode).
•
Agents retrieve but don’t condition on retrieval. In controlled interventions, withholding one decisive record cut action probability by 37.2–45.2 pp, but agents still acted in 46.5–53.5% of withheld episodes. Of 66 withheld-and-acted episodes, 65 included a call to the affected tool. They queried the thing, got no record, and acted anyway. A requester’s conflicting claim suppressed action more than silent absence of evidence.
•
If you evaluate agents today with endpoint checks or \u03c4-bench-style milestones, assume your success numbers overstate reliability. Specifically, assume your agent sometimes acts on evidence it never actually gathered, especially on multi-step workflows. Worth sampling trajectories and asking: before each state-changing call, is the justifying observation in the transcript, and does it refer to the correct entity?
•
The gap between Legacy (>95%) and interactive execution (often <50%) means a model that aces a function-call-selection eval can still ship an unsafe agent. Treat static tool-use benchmarks as a floor, not a predictor.
•
Harness choice has real effects, but they are model-specific and often redistribute failure modes rather than fix them. Worth testing the same model under at least two harnesses on your own workload before committing.
•
If you’re adding guardrails, this paper suggests explicit requester-side contradictions work better than passive absence of evidence or generic urgency warnings as a brake on premature action. Distractor records increase retrieval calls but don’t change action rates.
•
Artifacts are at safeact.github.io.
•
The evaluator scores observable support, not what the model internally relied on. An agent could technically have the information in context from an earlier call and still be flagged if provenance binding fails. The authors are explicit about this.
•
Domains are synthetic. Results don’t certify any real-world refund system, healthcare workflow, or infra tool.
•
The intervention studies run on 43 V1 cases with three configs. Directional, not definitive.
•
The supplementary SCGR method is framed as an analysis instrument, not a proposed fix. It improves some configs on some protocols and hurts others, so don’t read it as a recommended mitigation.