Research questionHow can mask-free video virtual try-on maintain garment consistency under motion, occlusion, and changing viewpoints?Mask-based localization can fail during large body motions or severe clothing occlusions. Sparse keyframes and limited multi-view data further make it difficult to preserve garment details consistently across frames and viewpoints, while video-level pseudo-data construction is expensive.