VideoMAE
PulseAugur coverage of VideoMAE — every cluster mentioning VideoMAE across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New Arti-JEPA model adapts video models for vocal tract MRI analysis
Researchers have developed Arti-JEPA, a new joint embedding predictive architecture designed to model real-time MRI data of the vocal tract for speech analysis. This model was trained on approximately 62 hours of unlabe…
-
Study finds DINOv2 effective for resource-limited self-supervised learning
A recent study explored self-supervised learning (SSL) for image and video pretraining under resource constraints, comparing various objectives. The research found that DINOv2-style pretraining performed best with limit…
-
New Sparse Autoencoders Enhance Video Representation Interpretability
Researchers have developed spatio-temporal sparse autoencoders (SAEs) to improve the interpretability and temporal coherence of video representations. Standard SAEs, while good at decomposing features, often sacrifice t…
-
AI models for heart health are spatially accurate but temporally blind
A new research paper published on arXiv investigates the attribution methods used to explain the decisions of deep learning models in echocardiography. The study found that while these models can accurately estimate lef…
-
New Transformer framework improves distracted driver detection
Researchers have developed a two-stage Transformer framework for accurately and efficiently localizing distracted driver behaviors in video streams. The framework combines VideoMAE for feature extraction with an Augment…
-
New datasets and models advance sign language recognition and translation
Researchers have developed new methods for sign language recognition and translation. One approach uses a deep learning pipeline combining a VideoMAE video transformer for classifying sign gestures into English words an…
-
New MoFore Framework Advances Self-Supervised Video Representation Learning
Researchers have introduced MoFore, a novel framework for self-supervised video representation learning that focuses on forecasting future latent embeddings from distant context clips. Unlike previous methods that relie…
-
Video foundation models show emergent intuitive physics understanding
A new research paper investigates whether video foundation models possess an understanding of intuitive physics. The study probes frozen representations of models like V-JEPA, VideoMAE, and LTX-Video using benchmarks su…
-
New AI models detect horse eye blinks for welfare assessment
Researchers have developed and evaluated three methods for automatically detecting and classifying horse eye blinks from video footage. These methods, including a YOLOv12 detector, an optical flow approach, and a fine-t…
-
New method steers physics reasoning in video world models
Researchers have developed a method called physics steering to control the physical reasoning of video world models. This technique uses a linear probe's weight vector, identified as a Concept Activation Vector (CAV), w…