Video-MME-v2
PulseAugur coverage of Video-MME-v2 — every cluster mentioning Video-MME-v2 across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New benchmarks and methods enhance audio-visual reasoning in LLMs · 2 sources tracked
Researchers have introduced new methods and benchmarks to improve audio-visual joint reasoning in omni-modal large language models. The OmniReasoning project developed OmniReasoningBench, a benchmark and data engine des…
-
Kwai releases Keye-VL-2.0 for long-video understanding
Kwai has released Keye-VL-2.0-30B-A3B, an open-source multimodal foundation model designed for long-video understanding and agentic intelligence. This model utilizes DeepSeek Sparse Attention to process up to 256K conte…
-
GridProbe cuts VLM compute cost for long videos
Researchers have developed GridProbe, a novel method to improve the efficiency of long-video Visual Language Models (VLMs). This technique adaptively selects relevant frames during inference, reducing the computational …