Raven is a multi-agent system where a Host Agent decomposes a user request into a typed DAG, assigns each node to a specialist model-harness pair (built-in or third-party), validates the graph before dispatch, and routes artifacts through recorded handoffs. The paper also gives sufficient conditions under which such composition reliably covers tasks no single agent can solve alone within a shared budget.
As you push LLM agents past single-turn tasks into long workflows (research + code + slides + an on-call monitor), two things break. First, the harness, meaning the tools, context manager, retry rules, and verifiers wrapped around the model, keeps growing. Hand-tuning it for every domain stops scaling. Second, specialists built independently (think Claude Code, Codex agent harness, and in-house agents) have incompatible interfaces, so stitching them together through free-form chat loses artifacts and context.
Typical baselines today either run one powerful agent on everything, or use conversation frameworks like AutoGen and MetaGPT where roles are configured but the plan is not an inspectable, typed object. Raven’s bet is that making the plan a validated DAG of model-harness pairs, with explicit artifact handoffs, is what lets heterogeneous agents actually cooperate.
A request flows through five stages. The Host Agent loads orchestration guidance on demand, builds a DAG where each node names an executor, a prompt template, dependencies, and input bindings. Before any worker runs, a runtime checks format, graph structure (acyclic, resolvable paths), executor capabilities (for example, is this agent stateful enough to share an instance handle), registration, and a shared dispatch budget. Then nodes execute as their predecessors settle; each result is checked by an independent completion judge. A failed node suspends and asks the host to continue with a message, abandon, or replan.
Artifacts are files, not paraphrases. Each node writes its prompt, output, and transcript to disk; successors read by path or inline. Identifiers persist across the whole conversation, so later graphs can depend on earlier nodes without re-executing them.
Four native specialists ride on this: Raven-Research (web + a page-reading reviewer), Raven-Code (reads repo instruction files, refuses edits to stale files, runs tests), Raven-Design (render-and-inspect for slides and SVGs), and Raven-Oncall (long-running jobs managed as persistent campaigns that wake the agent only when a measurement changes).
Two cross-cutting pieces: HarnessBank-style harness self-evolution searches patches to a frozen model’s harness, with paired-gain screening and a quality-diversity archive. EverOS handles memory (user profiles, episodes, agent cases) and feeds Skill Forge, which fuses local, memory, and SkillHub skills before an LLM gate picks at most two to inject.
# Host-mediated graph execution, simplified
plan = host.build_dag(request) # typed nodes + edges
if not runtime.admit(plan): return runtime.first_error()
while plan.has_pending():
for v in plan.ready_nodes(): # predecessors settled
out = dispatch(v, inputs=assemble(v, ledger))
verdict = judge(v, out)
if verdict.ok: ledger.write(v, out)
else: host.resolve(v, verdict) # continue | abandon | replan
return host.synthesize(ledger)
The theory formalizes each node as a contract (precondition, postcondition, transfer guarantee) with a local failure budget. If every provider realizes its contract and the host picks a compatible plan, success probability is at least (1 - planning_failure) * (1 - sum of local failures), clipped to zero.
The paper introduces MAOB, MAOB, 140 occupational requests paired with reference DAGs over four specialist domains (research, coding, content, oncall). Graphs are scored before workers run, so this measures planning, not execution. Compared to Claude Code and Hermes Agent on the same backbones (Qwen3.8-27B and DeepSeek-V4-Flash-0731), Raven leads on all four metrics (Node F1, Edge F1, partial-order accuracy, Exact Match) under both. Exact Match improves by +10.4 and +10.5 percentage points over the strongest baseline.
•
Research (DeepResearch Mixed, a blend of BrowseComp, FRAMES, Humanity’s Last Exam, and xBench-DeepSearch): Raven-Research is top on all three shared backbones, reaching 76.5% with DeepSeek-V4-Flash versus 68.9% for the strongest baseline. Five of six paired comparisons are significant by McNemar; one is not (p=0.21).
•
Code: wins six of eight benchmark-backbone cells. Biggest gap is on SWE-Refactor whole-repo migrations (+9.5 and +3.0 points). On SWE-bench Pro, 15 more tasks resolved out of 731 than Claude Code at the same backbone.
•
Design: on PresentBench scores 80.2 vs 78.3 (Claude Opus 5) and 72.9 vs 52.4 (GPT-5.6 Luna).
•
Oncall on the single-GPU autoresearch task: lower validation BPB than Claude Code at similar runtime, with lower reported token cost.
Reused from prior work: HarnessBank’s harness evolution improves held-out Pass@1 on all seven benchmarks with a frozen Qwen3.6-27B (largest gains +15.4 on AppWorld, +13.9 on BrowseComp+). SkillCorpus skills lift Raven more than they lift OpenClaw on two of three benchmarks.
Important caveat the authors flag: graph agreement is not task success. A good plan with a weak executor still fails, and benchmark access was not blocked, so absolute BrowseComp and HLE numbers may be inflated.
If you’re building a multi-specialist agent system, the concrete pattern worth borrowing is the typed-DAG admission check before any worker runs. Most failures the paper attributes to multi-agent systems (bad handoffs, missing credentials surfaced only mid-run, path references sent to an agent without local-file access) can be caught by five groups of static checks: format, graph structure, capability, status, environment. This is cheap and does not require adopting Raven.
If you specifically have independently built agents you want to coordinate, Raven’s execution-adapter approach (wrap each agent so it exposes a common invocation interface, keep its native tools and loop) is more realistic than rewriting them into one framework. The repo is at GitHub.
The harness self-evolution results are worth testing only if you can hold your task model frozen and run a diagnosis-then-patch loop with paired evaluation. The gains reported are on tasks withheld from evolution, but the authors note elsewhere that harness gains can transfer only partially across models.
For skill libraries, the finding worth taking seriously is that coverage predicts gain: tasks whose top retrieved skill scored above 0.75 improved by 25.1 points on SkillsBench, versus 2.2 points below 0.45. If you deploy a skill library, instrument coverage per task rather than averaging gain across your whole benchmark.
Do not read the MAOB numbers as evidence that Raven will deliver better end-to-end task quality. MAOB scores plans against a reference DAG; execution is measured separately per specialist.
The planning benchmark is constructed graph-first then back-translated into a request, filtered for obvious giveaways like “first, then, finally”. Paraphrased planning cues could survive this filter, and the reference graphs were authored by Claude Opus 5, so there is some risk of favoring systems that reason similarly.
Several headline results rely on prior work (HarnessBank evolution, SkillCorpus skills) and the paper does not fully re-specify every run setting. Reproducing those exact numbers requires the prior-work artifacts.
The theory establishes sufficient conditions for reliable composition and new-task coverage. It does not claim Raven’s actual implementation satisfies those conditions on any particular task, nor that composition beats a single well-equipped agent given the same tools.
Shared group memory, the layer where the host writes verdicts about each agent after a round, is disabled by default and not evaluated. Treat it as a design proposal, not a measured feature.
Commercial deep-research comparisons use provider-specific cost accounting; the same-backbone comparisons are the clearest evidence.