Get Started
Home
Topics
Search
Library
Research questionHow can zero-shot VLA transfer across unseen embodiments be assessed without confounding task or protocol differences?Performance differences across robot bodies can reflect changes in tasks, scenes, or evaluation protocols rather than embodiment transfer itself. This makes it difficult to determine what enables a VLA model trained on some robots to operate on a genuinely unseen one.
AI
Computer Vision
Evaluation & Benchmarks
Machine Learning
Multimodal Models
Robotics
Latest papersRecent research connected to this question, newest first.ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop ManipulationEvidence covers a controlled benchmark with 14 held-out target embodiments, simulation and real-world validation, and two-finger grippers. It distinguishes strict zero-shot transfer, where the target embodiment is absent from all training data, from pretrain-exposed zero-shot transfer, where it appears only during pretraining; the study analyzes state-action representations, source embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure. Results are limited to stationary tabletop manipulation and do not establish performance for mobile bases, dexterous hands, or long-horizon tasks.research paper · Sep 5, 2026
Related questions
How can driving VLAs transfer zero-shot across unseen datasets and camera rigs as training diversity grows?How can offline multi-agent policies transfer zero-shot to unseen tasks with different agent counts?How can we diagnose vision-language-action models’ failures on spatially ambiguous, long-horizon manipulation tasks?How can zero-shot voice conversion transfer an unseen speaker’s identity while preserving content in low-latency streaming?