ASTRA: Adaptive Spatio-Temporal Representation for Action Recognition via Importance-Guided Motion Learning
IEEE Journal of Selected Areas in Sensors — accepted, to appear — 게재 확정
Abstract초록
3D-CNN action recognizers learn appearance and motion in an entangled way, so they lean on static visual cues instead of true motion — which hurts generalization and interpretability on resource-constrained AIoT edge devices. ASTRA is a dual-pathway architecture that separates semantic appearance from pure motion through three components: an STFormer module that models space and time adaptively through cross-attention and estimates frame-wise importance; Importance-Guided Temporal Perturbation, which isolates appearance-free motion by comparing original and perturbed inputs; and a Motion Gated Unit that refines the motion features through dual-gate modulation. Under a leak-free, video-level protocol ASTRA reaches 97.88% on UCF-101 and 81.05% on HMDB-51 with an MViT-v2-S backbone. Against its own R(2+1)D-18 backbone, the motion pathway adds a significant +2.81%p on HMDB-51 (McNemar p < 0.001) while staying equivalent on appearance-saturated UCF-101, at 32.9M parameters, 41.9 GFLOPs and 8.2 ms GPU latency.