Research questionHow can sign language translation models capture asynchronous lip cues and recognize fingerspelled terms without detailed supervision?End-to-end systems must combine signing components that differ in content and timing. Limited detailed supervision can make proper nouns and technical terms difficult to recognize while leaving lip movements underused for disambiguation.