PulseAugur
EN
LIVE 20:43:05
ENTITY VideoMME

VideoMME

PulseAugur coverage of VideoMME — every cluster mentioning VideoMME across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
0
10 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
10 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
RECENT · PAGE 1/1 · 10 TOTAL
  1. TOOL · CL_135427 ·

    Goal-Driven Data Optimization speeds up multimodal AI training

    Researchers have developed a framework called Goal-Driven Data Optimization (GDO) to improve the efficiency of multimodal instruction tuning. GDO computes sample descriptors to create optimized training subsets tailored…

  2. RESEARCH · CL_128785 ·

    New methods tackle OmniLLM token compression for efficiency

    Two new research papers propose methods to compress token sequences in omnimodal large language models (OmniLLMs) to reduce inference costs. The first paper, DASH, uses audio cues to dynamically segment sequences and a …

  3. TOOL · CL_117509 ·

    New STAR Framework Boosts LLM Video Analysis Capabilities

    Researchers have developed a Spatiotemporal Reasoning Framework (STAR) to enhance the video question answering capabilities of multimodal large language models (MLLMs). STAR equips models like GPT-4o with a Video Toolki…

  4. RESEARCH · CL_117441 ·

    VisReflect framework improves LVLM fine-grained perception in long contexts

    Researchers have introduced VisReflect, a novel framework designed to enhance fine-grained perception in Large Vision Language Models (LVLMs) when processing high-resolution images and long videos. This method addresses…

  5. RESEARCH · CL_115210 ·

    Reflect-R1 framework improves AI video understanding with evidence-driven self-correction

    Researchers have introduced Reflect-R1, a novel framework designed to enhance self-correction in long video understanding models. This system addresses the issue of models becoming overconfident due to a lack of externa…

  6. RESEARCH · CL_97982 ·

    OmniAgent uses active perception for efficient video understanding · 2 sources tracked

    Researchers have introduced OmniAgent, a novel omni-modal agent designed for video understanding that utilizes an iterative Observation-Thought-Action cycle based on Partially Observable Markov Decision Processes (POMDP…

  7. TOOL · CL_51670 ·

    CREST method efficiently selects key frames from long videos

    Researchers have developed CREST, a novel method for efficiently selecting key frames from long videos. This training-free approach leverages the temporal geometry of query-frame relevance, specifically focusing on loca…

  8. TOOL · CL_30588 ·

    AdaFocus framework boosts long video understanding with adaptive sampling

    Researchers have developed AdaFocus, a new framework designed to improve the efficiency of understanding long videos. This method avoids the high costs of dense encoding or the information loss from aggressive compressi…

  9. RESEARCH · CL_15643 ·

    New AI methods enhance video reasoning by structuring and selecting visual evidence

    Researchers are developing new methods to improve how large vision-language models (VLMs) understand and reason about long videos. Several papers introduce techniques for more efficient frame selection and evidence gath…

  10. TOOL · CL_15615 ·

    VideoThinker framework improves lightweight MLLMs' video reasoning via causal debiasing

    Researchers have developed VideoThinker, a novel framework designed to enhance the reasoning capabilities of lightweight multimodal language models (MLLMs) in video analysis. This approach addresses the issue of percept…