PulseAugur
EN
LIVE 22:42:57
ENTITY Mscoco

Mscoco

PulseAugur coverage of Mscoco — every cluster mentioning Mscoco across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 6 TOTAL
  1. RESEARCH · CL_227216 ·

    New research optimizes visual token processing for long-video MLLMs

    Researchers are exploring methods to optimize how multimodal large language models (MLLMs) process visual information, particularly for long videos. Several papers introduce techniques for selecting, compressing, and pr…

  2. TOOL · CL_143793 ·

    Xray-Visual model scales vision tasks with 15B image-text pairs

    Researchers have introduced Xray-Visual, a novel vision model architecture designed for large-scale image and video understanding. Trained on a massive dataset of over 15 billion image-text pairs and 10 billion video-ha…

  3. TOOL · CL_53654 ·

    New FAST-GOAL method enhances vision-language models for detailed text

    Researchers have developed FAST-GOAL, an efficient fine-tuning method designed to improve the ability of vision-language models like CLIP to process lengthy and detailed text descriptions. The method employs two main co…

  4. RESEARCH · CL_53958 ·

    Google DeepMind unveils Gemini Embedding 2 multimodal model

    Google DeepMind has introduced Gemini Embedding 2, a new native multimodal embedding model. This model can generate unified representations for video, audio, image, and text data, demonstrating strong zero-shot capabili…

  5. RESEARCH · CL_18576 ·

    Researchers unveil new stealthy backdoor attacks on AI models using diffusion and style features

    Researchers have developed new methods for backdoor attacks on advanced AI models, specifically targeting Vision-Language Models (VLMs) and Diffusion Models (DMs). One approach, CBV, uses diffusion models to create natu…

  6. RESEARCH · CL_11442 ·

    Researchers find single hub text exploits vulnerabilities in CLIP cross-modal encoders

    Researchers have identified a vulnerability in cross-modal encoders like CLIP, which map text and images into a shared embedding space. They discovered that a single "hub text" can generate high similarity scores with n…