Researchers have developed CoDAT, a Collaborative Dual-Attention Transformer designed for efficient action recognition on edge devices. This model utilizes a lightweight dual-branch attention mechanism, combining Spatial Convolutional Attention (SCA) for local feature aggregation and Strided Single-Head Attention (SSHA) for global context, significantly reducing computational costs. CoDAT also incorporates a parameter-free TShift module for temporal modeling, enabling efficient communication across frames. Experiments show CoDAT achieves a superior energy-accuracy balance compared to existing methods on various benchmarks, offering faster throughput and fewer parameters for real-time applications in edge IoT systems. AI
IMPACT Enables more efficient and real-time AI-powered perception systems on resource-constrained edge devices.
RANK_REASON The item is a research paper detailing a new model architecture and its performance on benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]
- Collaborative Dual-Attention Transformer
- EfficientViT384
- FastViT-S12
- ImageNet-1K
- Internet-of-Things
- Kinetics-400
- LAPS
- MA-52
- Jetson AGX Orin
- Raspberry Pi 5
- Spatial Convolutional Attention
- Strided Single-Head Attention
- TokShift
- TShift module
- UniFormer-B
- ViT-Temporal-Shift
- VSwin-T
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →