MathVista
PulseAugur coverage of MathVista — every cluster mentioning MathVista across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New benchmark reveals AI's struggle to draw geometric diagrams
A new benchmark, "Solving Is Not Drawing," has been introduced to evaluate the distinct capability of foundation models to construct geometric diagrams, a skill separate from mathematical problem-solving. The benchmark …
-
New benchmarks and models tackle AI-generated scientific diagrams · 4 sources tracked
Researchers have developed new benchmarks and models to address the challenge of generating scientifically accurate diagrams using AI. Princigram, a new generator, utilizes a Structured Physical Chain-of-Thought (SP-CoT…
-
New SD-MAR framework boosts VLM analytical reasoning across multiple images
Researchers have introduced SD-MAR, a new framework designed to enhance the analytical reasoning capabilities of vision-language models (VLMs) across multiple images. This framework utilizes synthetic data generated thr…
-
New dataset and model enhance multimodal math reasoning with diverse perspectives
Researchers have introduced MathV-DP, a new dataset designed to improve multimodal mathematical reasoning by capturing diverse solution trajectories for each image-question pair. This dataset aims to provide richer supe…
-
Self-Improving VLMs Can Regress on New Tasks, Study Finds
A new research paper reveals that self-improving visual-language models (VLMs) can regress on new tasks, contrary to the assumption that stronger verifiers always yield stronger students. The study found that verifier q…
-
Research: Stage-1 training impacts VLM entropy, not final outcome
A new research paper explores the impact of different Stage-1 training methods on vision-language models (VLMs). The study found that while Stage-1 training, such as supervised fine-tuning (SFT) or on-policy distillatio…
-
UnAC method enhances LMMs for complex multimodal reasoning with adaptive prompting
Researchers have introduced UnAC, a novel multimodal prompting method designed to enhance the reasoning capabilities of Large Multimodal Models (LMMs) on complex visual tasks. This method employs adaptive visual prompti…
-
New CGC framework boosts multimodal LLMs for fine-grained image understanding
Researchers have introduced Compositional Grounded Contrast (CGC), a new framework designed to enhance the fine-grained multi-image understanding capabilities of Multimodal Large Language Models (MLLMs). This approach a…
-
OpenAI's new models let ChatGPT think with images for advanced reasoning
OpenAI has introduced its latest visual reasoning models, o3 and o4-mini, which allow AI to "think with images" as part of its internal reasoning process. These models can perform image manipulations like cropping and z…