Research questionHow can video models recognize unseen actions from only a few labeled examples?Few-shot action recognition must infer action categories from sparse annotated videos, where informative spatial and temporal cues can vary across frames and classes.