Get Started
Home
Topics
Search
Library
6 min read · Agents · Robotics · Added Oct 10 · Paper published Oct 8, 2026

SuperNav: An Agentic Navigation System for Any Task in Any Scene

Source: research paper via Hugging Face Daily Papers
SuperNav skips fine-tuning a vision-language-action model for robot navigation and instead wraps a frozen MLLM in a tool-use harness with a pixel-pointing interface and Markdown “Skills.” On unseen scenes it hits 78% success vs 34% for the strongest fine-tuned video-policy baseline.
TL;DR
SuperNav turns a general Multimodal Large Language Model into a navigation agent by giving it tools, a visual-point interface, and reusable Skills instead of fine-tuning it on navigation data, letting one system handle object, multi-object, and need-based requests across unseen scenes.
Why It Matters
If you want a home or service robot that can “find my keys,” then “visit the kitchen, then the office,” and then “find somewhere I can sit,” you basically have two options today, and both are painful.
One is a modular pipeline: a hand-wired state machine that uses a vision-language model to score how relevant each view is to the goal, then hands off to exploration and stopping logic. These work, but every new kind of request (instance vs. category vs. abstract need) means rewriting the workflow. The other is end-to-end fine-tuning: train a vision-language model on paired instructions and robot trajectories to directly output actions, like NaVid or UniNaVid. That gives you one model, but it only generalizes as far as the training trajectories cover, and the authors report large task-completion gaps when these models are run in scenes they weren’t trained on.
SuperNav’s bet is that a modern Multimodal Large Language Model already knows what a sofa looks like, what “a place to rest” means, and how to decide whether a candidate matches. What it lacks is a body. So give it one through tools, and leave the model itself untouched.
How It Works
Think of the Multimodal Large Language Model as the brain in an agent loop, similar to how a coding agent calls shell commands and reads their output. At each step it sees the current four-view RGB observation, the request, and the running history, and picks one tool call: look around, turn, move to a point, update which goals are done, read a Skill, or close the session.
The clever interface piece is the unified visual-point interface. Instead of outputting low-level actions (forward, turn left) or XY coordinates in some map frame, the model picks a pixel in an image: “go to point (0.42, 0.78) in the front view.” A motion backend turns that pixel into actual motion. Two backends are interchangeable: a Geo-based Executor that uses depth and a navigation mesh to plan a path, and a Learned Executor adapted from NoMaD that predicts short motion segments directly from the marked image and recent frames. Both return fresh four-view images plus execution feedback, so the model’s decision loop stays the same regardless of which backend is driving.
Navigation Skills are Markdown documents the model can load on demand. They describe reusable procedures for searching (record branches you haven’t explored), recovery (if a passage is blocked, back off and try another route), and completion (verify the candidate’s appearance and relations before declaring success). Skills are advice entering the context, not hardcoded control flow. The model still chooses every action.
Context management matters for long episodes. The harness keeps textual history, task progress, and image file paths, but strips the actual image bytes of old observations once newer ones supersede them. The model can re-request an older image when it wants it.
obs = init_scene() while not done: ctx = (request, goal_state, history, obs) action = mllm.decide(ctx, skills_available) if action.type == "point_nav": obs = point_nav(view=action.view, uv=action.uv) elif action.type == "read_skill": ctx += load_skill(action.name) elif action.type == "close": done = True # declares achieved or blocked
What They Found
SuperNav is evaluated across four task families, using GPT-5.6 Terra as the decision model.
•
On a 150-task single-object instance benchmark built in Habitat-GS with InteriorGS scenes, SuperNav reaches 78.00% success vs. 34.00% for the strongest video-policy baseline, UniNaVid. Success requires geodesic distance under 1 m to an approved viewpoint plus an explicit STOP.
•
On demand-driven navigation (200 tasks from Demand-Bench in AI2-THOR, where goals are inferred from needs like “prepare a workspace”), SuperNav hits 59.50% vs. 37.50% for OmniNav’s Action Former. Both motion backends beat the baselines.
•
On a 120-episode subset of HM3D-OVON for category-level navigation, with Geo-based Executor and no OVON-specific training, it reaches 73.33% SR / 0.4105 SPL at the 1 m threshold and 68.33% / 0.3837 at 0.25 m. Success weighted by Path Length measures how close the executed path is to the shortest reference path.
•
The ablations are the most informative part. Removing Navigation Skills hurts success rate more than restricting the model to a front-only view, especially at the stricter 0.25 m threshold. Replacing direct pixel pointing with an external language-grounding tool (built on LocateAnything) does not help. Swapping the decision model from Terra (high reasoning) to gpt-6-astra (medium) actually improves both SR and SPL with the harness fixed, which the authors frame as evidence the harness is compatible with different backing models rather than tuned to one.
•
A Unitree Go2 quadruped runs the same system with LiDAR-based geometric execution, recovering from blocked passages and verifying targets in real rooms.
Two honest observations about the numbers. The baselines were trained for action prediction, not called through tools, so some of the gap reflects paradigm mismatch rather than pure capability. And the published HM3D comparisons use different episode subsets and action budgets, which the paper flags explicitly.
What’s Useful
If you are building a robot stack and considering whether to fine-tune a vision-language-action model, this paper is a strong data point that a harness-plus-tools approach can match or beat fine-tuned baselines on cross-scene tasks without touching the model weights. Worth testing on your own tasks before committing to a fine-tuning pipeline.
The visual-point interface is independently reusable. Having the model output a pixel in a labeled view, and letting a separate backend turn that into motion, decouples perception-level decisions from locomotion. You can swap in whichever controller your platform supports (classical planner, learned policy, whatever drives your base) without rewriting prompts.
For agent builders outside robotics, the Skills-as-Markdown + media-only context pruning pattern is a clean recipe for long-horizon agents in any modality: let the model discover procedural guidance by name, load it on demand, and drop heavy payloads from history while keeping pointers so they can be re-fetched. The ablation showing Skills matter more than extra views is a useful prior.
One prerequisite: the geometric backend needs depth, calibration, and a navigation mesh or LiDAR map. If you only have RGB, you need the learned backend, which the paper shows retains much of the success but loses path efficiency.
Project page: zju3dv.github.io/SuperNav.
Caveats
Semantic judgments ride entirely on the underlying Multimodal Large Language Model. If it can’t tell a sofa from a loveseat, the harness won’t save you. The paper’s results use strong proprietary models, and the authors don’t report what happens with smaller open models.
The demand-driven evaluation scores ordered arrival at relevant objects, not whether the robot actually accomplished the activity (preparing the workspace, etc.). The reported SR is a navigation success, not a task-completion success.
Multimodal Large Language Model inference adds latency and variable cost per step, so episodes have unpredictable duration. For latency-sensitive deployments this is a real constraint. And the baseline comparisons on HM3D use different protocols than published numbers, so cross-paper leaderboard reading is not valid here. The authors are upfront about this.
Topics
Agents
Robotics
Agents
Robotics
Up next in Agents
From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation
OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents252 episodes
Robotics74 episodes