Researchers have developed Speech2MaskTrack, a novel approach for speech-guided referring video object segmentation. This method connects speech recognition with temporal grounding and mask tracking to identify objects based on spoken motion descriptions. Speech2MaskTrack achieved second place in the MeViS-Audio track of the 8th LSVOS Challenge by transcribing spoken queries into structured constraints and ranking instance tracks using motion and relation evidence. AI
IMPACT This research advances speech-guided video object segmentation, potentially improving human-AI interaction in video analysis tasks.
RANK_REASON The item is an academic paper detailing a runner-up solution for a challenge. [lever_c_demoted from research: ic=1 ai=1.0]
- generative pre-trained transformer
- LSVOS Challenge
- MeViS-Audio
- SAM3.1
- SaSaSa2VA
- Speech2MaskTrack
- TRACE
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →