Get Started
Home
Topics
Search
Library
Research questionHow can sign language translation models capture asynchronous lip cues and recognize fingerspelled terms without detailed supervision?End-to-end systems must combine signing components that differ in content and timing. Limited detailed supervision can make proper nouns and technical terms difficult to recognize while leaving lip movements underused for disambiguation.
AI
Computer Vision
Image & Video Processing
Multimodal Models
Natural Language Processing
Latest papersRecent research connected to this question, newest first.SignBind-LLM: Multi-Stage Modality Fusion for Sign Language TranslationThe evidence concerns a modular sign-language translation system with separate streams for continuous signing, fingerspelling, and lipreading, followed by temporal fusion and language decoding. The reported system uses CTC pretraining on approximately two million automatically generated pseudo-gloss sequences and is evaluated on How2Sign, BOBSL, and ChicagoFSWild+; the input does not establish broader language coverage or deployment behavior.research paper · Sep 2, 2026
Related questions
How can sign-language production models generate grammatical non-manual features without losing expressive variation?How can sign recognition from continuous signing identify unseen signs without gloss-annotated vocabularies?How can video-based sign dictionary retrieval generalize across signers and unseen languages?How can text-to-ASL gloss translation preserve discourse coherence across sentences?