Get Started
Home
Topics
Search
Library
Research questionHow can training-free air-ground navigation agents share aerial context to guide ground vehicles in closed-loop vision-and-language navigation?The UAV observes the environment from a global bird’s-eye view, while the ground vehicle must act from local first-person observations. Without reliable information sharing, the ground vehicle lacks the spatial context needed to navigate effectively.
AI
AI Agents
Computer Vision
Multi-agent Systems
Multimodal Models
Robotics
Latest papersRecent research connected to this question, newest first.Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye MapsThe evidence concerns an air-ground system with a UAV and UGV evaluated in 100 closed-loop CARLA-Air episodes in Town10HD. Reported results include a 77.0% joint success rate; the evidence is limited to this simulated setting.research paper · Sep 8, 2026
Related questions
How can closed-loop vision-language navigation learn effectively despite distribution shift and sparse micro-action rewards?How can embodied vision-language navigation follow instructions in unseen environments with limited data and memory?How can multimodal models control drones reliably under prompt-defined action protocols and terminate at the right time?How can vision-language models ground semantic driving inputs in physically plausible continuous actions with low latency?