ScienceQA
PulseAugur coverage of ScienceQA — every cluster mentioning ScienceQA across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New research explores "two clocks" in diffusion LMMs: answer stabilization vs. rationale unfolding
A new research paper introduces the concept of "two clocks" in diffusion Large Multimodal Models (LMMs), distinguishing between when an answer stabilizes and when its rationale is fully generated. The study analyzes thi…
-
New dataset SciGram boosts AI understanding of scientific diagrams
Researchers have developed a new framework to generate large-scale instruction data for scientific diagram understanding. This approach extracts concepts from scientific curricula, synthesizes facts, and retrieves relev…
-
New HANIA framework enhances multimodal question answering with graph-based evidence selection
Researchers have developed HANIA, a novel framework designed to improve multimodal question answering by using a planner-guided multimodal graph. This system extracts relevant visual and textual evidence, constructs a g…
-
New VQA Systems Enhance Document Understanding and Educational Reasoning
Researchers have developed two new approaches for multimodal visual question answering (VQA) systems. The first, Q-Guide, uses a small agent to intelligently acquire evidence by determining what information is missing a…
-
GraphLoom framework improves multimodal RAG with knowledge graphs
Researchers have introduced GraphLoom, a novel framework designed to enhance multimodal retrieval-augmented generation (RAG) systems. This system constructs a multimodal knowledge graph from various data sources, includ…
-
New TGIF module reduces hallucinations in multimodal LLMs
Researchers have developed TGIF (Text-Guided Inter-layer Fusion), a novel module designed to reduce hallucinations in multimodal large language models (MLLMs). Unlike previous methods that focus on text or static visual…
-
New VSSD technique enhances multimodal reasoning in smaller LLMs
Researchers have developed a new technique called Visual Saliency Steering Distillation (VSSD) to improve multimodal chain-of-thought (CoT) reasoning in smaller language models. VSSD uses attention maps from larger mode…
-
Visual token pruning impacts MLLM calibration, research finds
A new research paper investigates the impact of visual token pruning on the calibration of multimodal large language models (MLLMs). The study, published on arXiv, reveals that the method used for pruning tokens signifi…
-
New dataset trains LLMs for K-12 educational risk assessment · 3 sources tracked
Researchers have developed AIriskEval-edu-db2, a new dataset aimed at training and evaluating Large Language Models (LLMs) for assessing pedagogical risks in K-12 educational content. The dataset includes over 1,600 exp…