PulseAugur
EN
LIVE 21:12:55
ENTITY Video LLMs

Video LLMs

PulseAugur coverage of Video LLMs — every cluster mentioning Video LLMs across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
9
24 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
9
24 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/2 · 24 TOTAL
  1. TOOL · CL_185670 ·

    Huawei IJCAI 2026 papers highlight efficiency gains in AI models · 1 source tracked

    Huawei presented four papers at IJCAI-ECAI 2026, shifting focus from scaling model size to optimizing efficiency and design. One paper details a hierarchical Vision Transformer (ViT) scaled to 30 billion parameters, ach…

  2. TOOL · CL_185507 ·

    New SlotNarrative Interface Boosts Video-LLM Token Efficiency

    Researchers have introduced SlotNarrative, a novel interface designed to make Video Large Language Models (Video-LLMs) more token-efficient. This system organizes videos into persistent object narratives using compact o…

  3. RESEARCH · CL_180974 ·

    Two new methods tackle Video-LLM token compression for efficiency

    Two new research papers propose novel methods for compressing video tokens in Video Large Language Models (Video-LLMs) to improve efficiency. The first paper introduces NovaCov, a set-wise token compressor designed for …

  4. RESEARCH · CL_181051 ·

    New PhyCheck dataset evaluates Video LLMs' understanding of physical laws

    Researchers have introduced PhyCheck, a new dataset designed to evaluate and improve the physical law understanding capabilities of Video Large Language Models (VideoLLMs). The dataset includes coarse-grained and fine-g…

  5. TOOL · CL_154701 ·

    Video-HOCA benchmark reveals Video-LLMs struggle with anomaly explanation

    A new diagnostic benchmark called Video-HOCA has been introduced to evaluate the physical anomaly reasoning capabilities of Video Large Language Models (Video-LLMs). This benchmark utilizes an Ontological-Causal taxonom…

  6. TOOL · CL_145821 ·

    New benchmark diagnoses visual grounding in video LLMs

    A new paper introduces the Visual Dependency Gap (VDG) to assess the visual grounding capabilities of video large language models (LLMs). The VDG measures the difference in accuracy between models processing original vi…

  7. RESEARCH · CL_141071 ·

    New E-VQA Task Aims to Make Video LLMs More Transparent

    Researchers have introduced Evidence-Backed Video Question Answering (E-VQA), a new task designed to make Video Large Language Models (Video LLMs) more transparent. Current models often provide answers without clear vis…

  8. RESEARCH · CL_141275 ·

    New SLVMBench benchmark reveals video LLMs struggle with skill learning from long memory

    Researchers have introduced SLVMBench, a novel benchmark designed to evaluate the ability of video large language models (video-LLMs) to learn skills from extended video memory and apply them in real-time scenarios. The…

  9. RESEARCH · CL_139317 ·

    GeoTrace framework compresses video tokens for efficient Video LLMs

    Researchers have introduced GeoTrace, a novel framework designed to enhance the efficiency of Video Large Language Models (Video LLMs) by compressing visual tokens. This training-free method decomposes video evidence in…

  10. RESEARCH · CL_128644 ·

    TimeThink framework enhances temporal reasoning in Video LLMs · arXiv paper

    Researchers have introduced TimeThink, a novel reinforcement learning framework designed to enhance the temporal reasoning capabilities of Video Large Language Models (Video-LLMs). This approach focuses on optimizing th…

  11. RESEARCH · CL_111633 ·

    Denoising Attention (DnA) improves visual task performance

    Researchers have introduced Denoising Attention (DnA), a novel method designed to improve the performance of attention-based models in visual tasks. DnA addresses the issue of noisy attention patterns produced by standa…

  12. RESEARCH · CL_79694 ·

    New benchmarks and frameworks enhance video temporal grounding

    Researchers have introduced new benchmarks and frameworks for improving temporal grounding in long-form videos. One study posits that hour-scale video grounding is primarily a search problem, not a recognition one, and …

  13. TOOL · CL_77289 ·

    New MACD method combats video LLM hallucinations

    Researchers have developed a new inference strategy called Model-Aware Contrastive Decoding (MACD) to combat hallucinations in video language models. MACD leverages the model's own feedback to identify and target specif…

  14. TOOL · CL_66155 ·

    New framework measures video-LLM complexity using attribute analysis

    Researchers have introduced VideoABC, a new framework designed to measure the complexity of video-question pairs for video-LLMs. This non-parametric measure utilizes a vocabulary of video attributes, such as scene compl…

  15. TOOL · CL_65487 ·

    V-LynX framework integrates new modalities into Video LLMs

    Researchers have developed V-LynX, a framework that allows new modalities to be integrated into Video Large Language Models (LLMs) by leveraging an existing token interface. This method uses a lightweight auxiliary path…

  16. TOOL · CL_51673 ·

    LiteFrame boosts Video LLM frame scaling and cuts latency

    Researchers have developed LiteFrame, an efficient vision encoder designed to improve the performance of Video Large Language Models (Video LLMs) when processing extended video content. This new framework uses Compresse…

  17. TOOL · CL_45039 ·

    New CRPO method enhances video LLM spatiotemporal sensitivity

    Researchers have developed a new framework called Counterfactual Relational Policy Optimization (CRPO) to improve the spatiotemporal sensitivity of video large language models (Video LLMs). This method addresses the iss…

  18. RESEARCH · CL_44056 ·

    Video-LLMs suffer from directional motion blindness, researchers find

    Researchers have identified a significant limitation in current Video Large Language Models (Video-LLMs), termed "directional motion blindness," where models struggle to accurately perceive and articulate the direction …

  19. RESEARCH · CL_47629 ·

    New frameworks and benchmarks advance Video-LLM efficiency and understanding

    Researchers have introduced EarlyTom, a novel framework designed to enhance the efficiency of video large language models (Video-LLMs) by compressing visual tokens early in the vision encoder. This approach significantl…

  20. TOOL · CL_25592 ·

    Video-LLMs struggle with temporal information flow, researchers find

    Researchers have identified a significant bottleneck in how Video Large Language Models (Video-LLMs) process temporal information, hindering their ability to understand the direction of video playback. While video-centr…