VQA v2
PulseAugur coverage of VQA v2 — every cluster mentioning VQA v2 across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New VQA research explores answerability prediction, counterfactual learning, visual benchmarks, and privacy
Researchers are advancing Visual Question Answering (VQA) through several new approaches. One paper introduces VT-Transformer, which uses a Transformer architecture to predict answerability by analyzing visual and textu…
-
New frameworks enhance knowledge-based visual question answering systems · 5 sources tracked
Researchers are developing advanced frameworks to improve Knowledge-based Visual Question Answering (KB-VQA) systems. These new methods focus on enhancing the retrieval of relevant external knowledge and ensuring that t…
-
New ASR method prevents multimodal LLMs from forgetting skills
Researchers have introduced Attention-Spectrum Regularization (ASR), a novel framework designed to prevent multimodal large language models (MLLMs) from forgetting previously learned skills when adapting to new data. AS…
-
ITNet architecture unifies convolution, attention, and recurrence
Researchers have introduced ITNet, a novel neural network architecture that unifies convolution, attention, and recurrence into a single learnable integral transform. This architecture uses a learnable kernel, implement…
-
New research tackles continual learning in LLMs with novel MoE methods
Two new research papers propose novel approaches to continual learning in large language and vision-language models, aiming to mitigate catastrophic forgetting. CP-MoE introduces a transient expert to guide updates and …
-
Researchers unveil new stealthy backdoor attacks on AI models using diffusion and style features
Researchers have developed new methods for backdoor attacks on advanced AI models, specifically targeting Vision-Language Models (VLMs) and Diffusion Models (DMs). One approach, CBV, uses diffusion models to create natu…