visual perception
PulseAugur coverage of visual perception — every cluster mentioning visual perception across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
JEPA-Anything framework enables domain-agnostic world modeling
Researchers have introduced JEPA-Anything, a new framework designed for domain-agnostic world modeling. This approach extends joint-embedding predictive architectures by employing orthogonal predictive factorization (OP…
-
New benchmark Tri-PvP reveals modality bias in omni-modal LLMs
Researchers have developed Tri-PvP, a new benchmark designed to expose modality bias in omni-modal large language models (OLLMs). This benchmark addresses a limitation in previous evaluations by separating perceptual si…
-
New ML models tackle prediction of rare, large-scale events
Researchers have developed new machine learning models, including a Fourier-Mellin Neural Operator and a wavelet-decomposition based Graph Neural Network, to address the challenge of predicting rare, large-scale events …
-
Diffusion models show inherent attention mechanisms and improved sampling techniques · 7 sources tracked
Recent research explores advancements in diffusion models, a dominant architecture for image generation. One paper reveals that these models inherently utilize an attention mechanism similar to transformers, suggesting …
-
New theory explains AI model representations across modalities
Researchers have developed a new theory that explains the internal workings of AI models across different modalities like vision, audio, and language. This theory posits that classification tasks create a shared represe…
-
New TB-CSPN architecture enables safe UAV/UGV swarm coordination
Researchers have developed a new coordination architecture called TB-CSPN for heterogeneous UAV/UGV swarms. This system synthesizes mission actions from multi-modal sensor data, including radar, RF, acoustic, and visual…
-
Study finds multimodal LLMs perpetuate gender bias in musical instrument associations
A new study published on arXiv investigates gender bias in multimodal large language models (LLMs) by examining their associations with musical instruments. Researchers developed the Symphony-Bias dataset, which include…
-
DAP-Pose achieves state-of-the-art in robust multi-modal pose estimation
Researchers have developed DAP-Pose, a novel end-to-end model for robust multi-modal pose estimation. This system integrates visual, inertial, and GNSS measurements using a Bi-level Cross-modal Fusion (BCF) module to ca…
-
Li Auto aims for Tesla FSD V14 parity with self-developed AI chips and models
Li Auto is developing its autonomous driving capabilities to match Tesla's FSD V14, focusing on safety, efficiency, and comfort, alongside advanced features like recognizing special vehicles and traffic police signals. …
-
New ReTeX framework recovers task expert performance from merged AI models
Researchers have developed a new framework called ReTeX to address parameter interference in multi-task model merging. This method models interference as additive offsets and predicts these offsets to recover individual…
-
NVIDIA Vera CPUs to accelerate agentic AI for science at Los Alamos National Laboratory
NVIDIA's Vera CPU is set to power new supercomputers at Los Alamos National Laboratory (LANL), enhancing scientific discovery and agentic AI capabilities. The new systems, named Mission, Vision, and Veritas, will integr…
-
Apple Inc. plans dense product launches for 2026-2027, including smart glasses and AI-enhanced AirPods
Mark Gurman reports that Apple Inc. is planning an exceptionally dense product release schedule for 2026 and 2027. Key launches include new iPhone and Apple Watch models in Fall 2026, alongside Mac updates and new iPads…
-
DeepSeek unveils new vision model for enhanced perception
DeepSeek has announced the release of its new vision model, designed to enhance visual perception capabilities. The model is accessible via a provided chat link, signaling a step forward in the company's AI development.
-
AI Co-Scientist automates research loop, boosts search ranking performance
Researchers have developed an AI Co-Scientist framework that integrates LLM agents with direct cloud-compute access to automate the research loop for search ranking systems. This framework utilizes a hybrid agent archit…
-
PaperJury system streamlines LaTeX paper review with deterministic orchestration
Researchers have developed PaperJury, a novel system designed for the rigorous review and revision of LaTeX documents, particularly for scientific papers. This closed-loop system separates deterministic orchestration fr…
-
New Graph Learning Frameworks Tackle LLM Noise and Heterogeneity
Researchers are developing new methods for graph learning that leverage or bypass large language models (LLMs). One approach, CANE, addresses the issue of noisy LLM-generated labels by estimating cluster-conditional rel…
-
EU governments ditch US tech for local, open-source alternatives
European governments are increasingly shifting away from US technology giants towards local and open-source alternatives. This strategic pivot is driven by concerns over data sovereignty, control, and reducing dependenc…
-
New method enables generalizable neural scaling laws across domains
Researchers have developed a method to create generalizable neural scaling laws that can be applied across different domains. These laws predict the relationship between model performance and resources like data or comp…
-
New tool cuts GPU memory use in AI training by optimizing optimizer states
Researchers have developed a Budget-Aware Optimizer Configurator (BAOC) to address the significant GPU memory consumption during large-scale model training. BAOC intelligently assigns different optimizer configurations …
-
New SIGHT framework improves wearable activity recognition with temporal structure
Researchers have developed SIGHT, a new test-time adaptation framework designed to improve the performance of wearable human activity recognition models. This framework addresses performance degradation caused by shifts…