Research questionHow can ViT-based video facial-expression recognition detect subtle, localized temporal changes that global attention overlooks?Subtle expression cues may appear in small facial regions and persist for only short intervals. Video Transformers can instead emphasize dominant motion and coarse temporal patterns, making those fine-grained changes difficult to recognize.