Q-Former
PulseAugur coverage of Q-Former — every cluster mentioning Q-Former across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New gait retrieval task and dataset introduced by researchers
Researchers have introduced Composed Gait Retrieval (CoGR), a new task focused on retrieving gait sequences based on a reference sequence and a natural language modification query. To support this task, they developed t…
-
New research enhances AI for clinically faithful medical image captioning · 2 sources tracked
Two new research papers explore advancements in medical image captioning, focusing on improving clinical faithfulness and accuracy. The first paper introduces a framework that enhances alignment between visual and textu…
-
New EEG foundation model EEG-PRIME improves cross-dataset decoding
Researchers have developed EEG-PRIME, a novel two-stage foundation model designed to improve electroencephalography (EEG) decoding across different datasets and subjects. This model combines masked pretraining with prot…
-
New framework learns implicit music styles for symbolic generation
Researchers have developed a novel cross-modal framework to learn and apply implicit music styles for symbolic music generation. The model, inspired by BLIP-2, utilizes a Querying Transformer (Q-Former) to extract style…
-
New LLM generates interpretable behavior descriptions for autonomous vehicles
Researchers have developed CommandLM, a novel multimodal large language model designed to generate human-readable descriptions of ego vehicle behavior from fused sensor data. This model integrates LiDAR and multi-camera…
-
New MEUSLI projector enables multilingual ASR and speech understanding
Researchers have developed MEUSLI, a novel multilingual projector designed to link speech encoders with large language models (LLMs) for advanced speech processing tasks. This system extends existing monolingual project…
-
New CAPTAIN method uses language models for APT detection with less data curation
Researchers have developed CAPTAIN, a new method for detecting Advanced Persistent Threats (APTs) in large-scale logs. Unlike previous approaches that require extensive data curation and preprocessing, CAPTAIN utilizes …
-
Video-LLMs struggle with temporal information flow, researchers find
Researchers have identified a significant bottleneck in how Video Large Language Models (Video-LLMs) process temporal information, hindering their ability to understand the direction of video playback. While video-centr…
-
CSMCIR framework enhances composed image retrieval with symmetric alignment
Researchers have introduced CSMCIR, a novel framework designed to improve composed image retrieval (CIR) by addressing the fragmentation of representation spaces in existing methods. This approach utilizes a Multi-level…
-
ViBE framework maps visual stimuli to M/EEG brain signals
Researchers have developed ViBE, a new framework for brain encoding that translates visual stimuli into magnetoencephalography (MEG) and electroencephalography (EEG) signals. The system utilizes a spatio-temporal convol…