Large Vision-Language Model
PulseAugur coverage of Large Vision-Language Model — every cluster mentioning Large Vision-Language Model across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New methods enhance visual document question answering with adaptive retrieval and agentic restoration
Researchers have developed new methods to improve Visual Document Question Answering (DocVQA) and Knowledge-Based Visual Question Answering (KB-VQA). ViSAR introduces an adaptive retrieval technique that dynamically sel…
-
New UniME-R1 framework improves multimodal retrieval with feedback-driven reasoning · 2 sources tracked
Researchers have developed UniME-R1, a novel framework designed to enhance unified multimodal retrieval by incorporating retrieval feedback into the reasoning process. Unlike previous methods that relied solely on query…
-
TextGaze uses LVLM for gaze target estimation · arXiv cs.CV
Researchers have introduced TextGaze, a novel architecture for gaze target estimation that utilizes a Large Vision-Language Model (LVLM) for semantic guidance. This approach aims to overcome the limitations of existing …