Researchers have developed SMART, a novel framework that leverages multimodal large language models (MLLMs) to improve sign language recognition and spotting. This approach uses MLLM-generated motion descriptions as auxiliary semantic cues and employs a stable video-text alignment method suitable for small-batch training. The framework also incorporates a Multi-Scale Temporal Adapter for enhanced temporal representation learning and a CSLR-guided spotting module called CSFormer for dense temporal localization. Experiments on four benchmark datasets demonstrate SMART's effectiveness in both recognition and spotting tasks. AI
IMPACT This framework could improve accessibility for deaf and hard-of-hearing individuals by enhancing the accuracy and efficiency of sign language interpretation technologies.
RANK_REASON The cluster contains a research paper detailing a new framework for sign language recognition and spotting. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CSFormer
- CSL-Daily
- Disaster and Safety KSL
- Hugging Face
- Large-scale KSL
- multimodal large language model
- PHOENIX14-T
- SMART
- Transformer++
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →