Get Started
Home
Topics
Search
Library
6 min read · Agents · Multimodal · Added Oct 1 · Paper published Sep 25, 2026

Omni-IO Skills: Harnessing Your Agent Omni-Native

Source: research paper via Hugging Face Daily Papers
0:00 / 8:03
Multimodal agents today either need retraining to ingest audio/video/3D or get bolted to specialist tools via prompts that fumble routing and reuse. Omni-IO Skills adds a scheduler layer: declarative skills, a dependency graph, and an asset registry lift input coverage from 40% to 100% on two host agents, no retraining.
TL;DR
Omni-IO Skills is a plug-and-play Agent harness that wraps an existing agent with hierarchical Agent Skills, a dependency graph scheduler, and a persistent asset store, letting hosts like GPT-5.6 Sol and Claude Sonnet 5 accept and produce audio, video, 3D, and documents without retraining the model.
Why It Matters
Suppose you want your agent to take a product photo and some notes and return a poster, a promotional video with sound effects, and a landing page. Today’s general-purpose coding agents (the paper names Codex and Claude Code as the reference point) are strong at planning and running code, but their output surface is mostly text and software. Audio, video, 3D, and polished documents sit behind external specialist models.
Two common fixes both fall short. Training an Omni foundation model that natively handles every modality means every new codec or decoder forces a model update. Hanging specialist tools off an agent via MCP (Model Context Protocol) gives it reach, but leaves the hard parts to the prompt: picking the right tool, passing files between steps, running independent steps in parallel, recovering from failures, and reusing an artifact next turn when the user says “redo just the landing page.” The paper’s claim is that this coordination problem deserves its own layer, separate from both the model and the tool.
How It Works
The harness sits between the host agent and a pool of external multimodal services, and it has four layers.
First, procedural knowledge is packaged as Agent Skills at three granularities. Atomic Skills are single operations (generate an image, understand a video). Expert Skills package a complete deliverable workflow such as poster design, including assembly and quality check. Scenario Skills sit at the application level (“Event Material,” “Job Application”) and decide which deliverables a vague request implies, then call the right Expert or Atomic Skills. Each Skill is a declarative record of applicability conditions, inputs, procedure, outputs, and relationships to other Skills, so the host agent can inspect and compose them without loading implementation code.
Second, when the host picks a Skill, it recursively expands the higher levels until every remaining step is directly executable. Those terminal steps become nodes in a Declare Execution Graph (DEG), where edges encode both control order and data flow. Independent nodes in the same “Wave” run concurrently; an image-conditioned video node waits for the image. The graph is validated for cycles and dangling references before anything executes; a failed node cancels its descendants but independent branches keep running.
Third, an MCP (Model Context Protocol) tool service maps each node’s task type to a concrete tool, and a Provider and Configuration layer binds that tool to a specific model, credentials, and fallback. Swapping a video provider does not touch Skill definitions.
Fourth, every output lands in a persistent Asset Registry with a stable ID, type, provenance link to its source asset, and the turn it was created in. Downstream nodes resolve inputs by asset ID rather than file path, and later turns can refer back to prior assets by reference, which get materialized as already-completed nodes in the new graph.
def run(request, host_agent): skills = select_skills(request) # Scenario / Expert / Atomic tasks = expand_to_atomic(skills) # recursive expansion deg = build_graph(tasks) # nodes + depends_on edges validate(deg) # acyclic, refs resolve while pending(deg): wave = [n for n in pending(deg) if ready(n)] results = parallel_execute(wave) # via MCP or host-native for r in results: asset_registry.register(r) # append-only, file-locked return assemble_deliverables(asset_registry)
What They Found
Evaluation uses UniM-90, with GPT-5.6 Sol and Claude Sonnet 5 as two separate host agents. Each host is tested twice: as-is, and with Omni-IO Skills loaded into the same environment with everything else held fixed. The paper explicitly says it is measuring within-agent lift, not comparing the two models.
The headline number is modality coverage, reported as input-support rate \u03c4, the share of instances where the agent can even ingest all required input modalities. Base GPT-5.6 Sol handles 40.00%; base Claude Sonnet 5 handles 38.89%. With the harness, both reach 100%. Audio, video, documents, and 3D inputs that the bare agents could not accept become usable.
On Semantic\u2013Quality Coupled Score (SQCS), a joint semantic-correctness and generation-quality metric, the relative variant (which multiplies the absolute score by coverage, so it reflects performance over the full test set) rises from 26.99 to 74.94 for GPT-5.6 Sol and from 27.82 to 77.78 for Claude Sonnet 5. The paper frames these as 47.95 and 49.96 point gains. Notably, even the absolute SQCS, measured only on the subset each base agent could handle, improves (67.49 \u2192 74.94 and 71.53 \u2192 77.78), so the harness is not merely trading coverage for quality on easy inputs. Strict Structure Score, which measures whether outputs match the requested structural layout, reaches 100.00 and 99.78.
Two caveats on reading these numbers. The comparison is harness-on vs harness-off for the same host; there is no ablation isolating which of the four layers (hierarchical Skills, DEG scheduling, MCP routing, Asset Registry) drives the gains. And UniM-90 is a 90-instance subset of a benchmark co-authored by this group, selected independently of agent capabilities per the paper, but still a specific slice.
What’s Useful
If you are building a multimodal agent product and currently gluing specialist APIs together with ad-hoc prompts, the design worth borrowing is the separation between declarative Skill definitions and a graph scheduler with an asset registry. The lift from putting independent generation steps in the same Wave compounds with the ability to re-run just the changed node on a follow-up turn.
The three-level Skill hierarchy is worth examining before adopting wholesale. The win is that a vague user request (“make event materials”) maps to a Scenario Skill that decides the deliverable set, rather than forcing the base model to invent that plan every call. If your users send precise single-artifact requests, you may only need the Atomic and Expert levels.
For evaluation, note what the paper does not establish. It does not isolate which layer causes the gains, does not compare against other agent harnesses or orchestration frameworks, and reports no latency or cost numbers. If you care about those, worth testing in your own setup. The Asset Registry’s append-only JSON with file-lock serialization is a reasonable starting point for cross-turn reuse but is not benchmarked for concurrent-session scale.
No code or dataset release is mentioned in the provided text.
Caveats
The arXiv ID (2609.31847) and the model names (“GPT-5.6 Sol,” “Claude Sonnet 5”) suggest this paper is either forward-dated or uses placeholder identifiers; treat the specific model-version claims accordingly. The evaluation is one benchmark subset (UniM-90) built on the authors’ own UniM benchmark, with no third-party baseline harness for comparison, so the result shows the harness closes modality gaps for two specific hosts but does not position the design against alternative orchestration approaches. There is no component ablation, so attribution of the gain to any single layer is not supported by the evidence. The paper also acknowledges that absolute SQCS gains and coverage gains cannot be cleanly separated: with full coverage, harder instances enter the denominator.
Topics
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents191 episodes
Multimodal112 episodes
NLP94 episodes