Researchers have introduced a new approach to detecting when someone in an egocentric video is speaking directly to the camera wearer. This method reframes the problem from a binary classification to an online, frame-level prediction, accounting for various non-talk-to-me (TTM) states like talking to others or self-talking. A new dataset, the Online TTM Dataset, was created with approximately 900,000 annotated frames to support this formulation. A multimodal model integrating audio, visual, and speech-semantic cues achieved a 75.5% frame-level F1 score on TTM detection. AI
IMPACT Improves understanding of social interactions in egocentric video, potentially enabling more sophisticated AI assistants.
RANK_REASON The cluster contains an academic paper detailing a new method and dataset for a computer vision task. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- Egocentric Videos
- Hugging Face
- Online TTM Dataset
- Talk-to-Me Detection
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →