Ego4D: Around the World in 3,000 Hours of Egocentric Video
PulseAugur coverage of Ego4D: Around the World in 3,000 Hours of Egocentric Video — every cluster mentioning Ego4D: Around the World in 3,000 Hours of Egocentric Video across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
Audio-first triage slashes VLM calls for egocentric video captioning
Researchers have developed a novel audio-first approach for efficiently captioning long egocentric videos. This method prioritizes audio cues to decide which video segments are most relevant for analysis by a vision-lan…
-
Robots learning human actions: Four approaches to bridge video data and robot control · 1 source tracked
A joint paper from Tsinghua University, Hong Kong University of Science and Technology, and Microsoft Research Asia proposes a unified framework for enabling robots to learn human actions from human videos. The research…
-
New SmartRes framework boosts egocentric visual grounding efficiency
Researchers have developed SmartRes, a novel framework designed to improve the efficiency of egocentric visual grounding by dynamically routing high-resolution image patches. This method addresses the high computational…
-
SmartRes framework boosts egocentric visual grounding efficiency
Researchers have developed SmartRes, a novel framework designed to improve the efficiency of egocentric visual grounding tasks. This method optimizes processing in the pixel space by dynamically routing high-resolution …
-
EgoPlay system enables event-triggered video editing for egocentric streams
Researchers have developed EgoPlay, a novel system for event-triggered video editing in egocentric streams. This system, fine-tuned from a V2V diffusion transformer using data primarily from Ego4D, can identify specific…
-
Future context improves gaze estimation, but only up to a point
Researchers have developed a framework to study the impact of future video frames on egocentric gaze estimation models. Their findings indicate that while future context improves causal gaze prediction, the benefits pla…
-
New dataset and model advance scene graph reasoning for human activity understanding
Researchers have introduced SG-Ego, a new dataset that extends Ego4D with spatio-temporal scene graphs to better understand human activities in first-person videos. They also developed GLEN, a graph-based model designed…
-
New benchmark LongEgoRefer challenges AI with long-form egocentric video comprehension
Researchers have introduced LongEgoRefer, a new benchmark designed to evaluate video referring expression comprehension in long-form egocentric videos. This benchmark, derived from the Ego4D dataset, features nearly 1,5…
-
FlexLAM introduces variable-length latent actions to improve video-based decision-making
Researchers have introduced FlexLAM, a novel approach to latent action learning that addresses the bottleneck trade-off in existing models. Unlike previous methods that use a fixed-capacity bottleneck, FlexLAM employs v…
-
New benchmark and architectures for proactive AI assistants released
Researchers have introduced EgoProactive, a new dataset and benchmark suite called Pro extsuperscript{2}Bench, designed to evaluate proactive procedural assistance systems. These systems aim to provide real-time, step-b…
-
New method fuses hand trajectory for egocentric video query grounding
Researchers have developed a new method for grounding natural language queries in egocentric videos by incorporating hand trajectory data. This approach fuses hand kinematic features with pre-trained video-text features…
-
FROST-STA system predicts object interactions in egocentric video
Researchers have developed FROST-STA, a system designed for short-term anticipation in egocentric videos, aiming to predict object interactions. The model uses frozen dense features from a ViT-G backbone, extracting vid…
-
New model advances behavioral recognition from AR glasses sensors
Researchers have developed a new method for recognizing complex human behaviors using data from head-mounted Inertial Measurement Units (IMUs), commonly found in AR smart glasses. They created a large dataset and a hier…
-
VISTA system wins Ego4D challenge with object interaction anticipation
Researchers have developed VISTA, a novel system designed for anticipating human-object interactions in egocentric videos. VISTA integrates spatial object detection with temporal context from a frozen V-JEPA 2.1 model t…