GPT-5.2
PulseAugur coverage of GPT-5.2 — every cluster mentioning GPT-5.2 across labs, papers, and developer communities, ranked by signal.
- developed by OpenAI 100%
- subsidiary of OpenAI 100%
- instance of LLM 90%
- instance of LLMs 90%
- instance of ChatGPT 90%
- instance of DeepSeek-V3.1 90%
- competes with Gemini-3.1 Pro 80%
- authored by arXiv 70%
- competes with GPT-4o 70%
- competes with Claude Sonnet 4.5 70%
- competes with Claude Opus-4.6 70%
- instance of Gemini-3.1 Pro 70%
7 day(s) with sentiment data
-
Open-Source vs. Proprietary LLMs: A Strategic Decision Framework · 3 sources tracked
The debate between open-source and proprietary Large Language Models (LLMs) is evolving, with open-source models increasingly closing the capability gap with their proprietary counterparts. While proprietary models like…
-
New benchmark reveals MLLMs struggle with egocentric puzzle assistance
Researchers have developed PuzzleMate, a new framework and benchmark designed to evaluate the capabilities of Multimodal Large Language Models (MLLMs) in providing step-by-step guidance for complex physical tasks, using…
-
LLMs struggle with deictic ambiguity in draft-verify-revise pipelines
A new research paper explores how Large Language Models (LLMs) in draft-verify-revise pipelines can struggle with deictic ambiguity, where context-dependent expressions like "previous" can refer to different things acro…
-
New framework improves funder name disambiguation in research publications
Researchers have developed a new framework for disambiguating funder names in scientific publication records, addressing challenges like spelling variations and abbreviations. By integrating datasets from the Research O…
-
New taxonomy reveals how AI models handle user disagreement
A new research paper introduces a taxonomy for understanding how large language models (LLMs) manage their epistemic authority, or claim to knowledge, when faced with user disagreement. The study analyzed over 32,000 re…
-
Human forecasters narrowly beat AI bots in FutureEval, but the gap is closing
In the latest FutureEval spring results, human forecasters narrowly outperformed AI bots, though the difference was not statistically significant. The AI bot team showed notable improvement over the past year, significa…
-
New local agent SciLENS synthesizes scientific literature, rivals GPT-5.2
Researchers have developed SciLENS, a novel autonomous agent framework for local scientific literature synthesis that operates without reliance on proprietary online services. This system integrates structural visualiza…
-
New features help detect and guide LLM-generated Korean poetry
Researchers have developed a method to detect and guide the generation of Korean poetry by large language models (LLMs). The approach uses interpretable form-level linguistic features, such as output length, line-final …
-
LLM reasoning exhibits irrationality beyond value alignment, study finds
A new research paper from arXiv explores the concept of "rational value risk" in large language models, suggesting that even well-aligned models can exhibit irrationality during reasoning. This risk is quantified as a d…
-
Chinese AI Models GLM & MiniMax Accessible Globally With Caveats · 1 source tracked
As of August 2026, Chinese AI models GLM (from Zhipu AI) and MiniMax are accessible outside China, though direct access presents challenges. Zhipu AI's international API is priced approximately double its domestic rate,…
-
New benchmark tests LLMs' ability to clarify search queries
Researchers have developed a new benchmark called Clarify-Then-Search to evaluate the effectiveness of Large Language Models (LLMs) in improving deep search capabilities. This benchmark, built on real-world query data, …
-
Legal AI models exhibit 'inertia of confidence', study finds
A new research paper explores the 'inertia of confidence' in legal AI models, where systems like ChatGPT, Meta AI, and Perplexity AI provide incorrect legal verdicts with high certainty. The study found Meta AI had the …
-
New research explores advanced jailbreak techniques and detection methods for LLMs and VLMs
Researchers are developing advanced methods to test the safety and robustness of large language and vision-language models against jailbreaking attempts. New frameworks like SEAV focus on validating the correctness and …
-
New VLM VFIG converts raster images to complex SVG diagrams
Researchers have developed VFIG, a new vision-language model (VLM) designed to convert rasterized images into Scalable Vector Graphics (SVG) format. This advancement addresses the common issue of lost original vector fi…
-
AI safety experiment tests LLM memory instruction adherence with larger user profiles
An AI safety experiment explored how the size of a user profile affects an LLM's adherence to instructions regarding memory usage. The study partially reproduced findings from the PersistBench paper, which investigates …
-
New benchmark SciFigBench tests VLM reliability with misleading scientific figures
A new benchmark called SciFigBench has been developed to evaluate vision-language models (VLMs) on their behavioral reliability when presented with incomplete or misleading scientific figures. The benchmark assesses per…
-
New SocialRL method trains small LLMs to match GPT-4/5 negotiation skills
A new research paper introduces SocialRL, a method to enhance the social reasoning capabilities of small language models (4B parameters). The SocialRL framework trains models to act as strategic negotiators rather than …
-
OpenAI expands inference residency to UAE, third market after US and EU
OpenAI has expanded its inference residency service to the United Arab Emirates, making it the third market, after the US and EU, to offer this capability. This new service allows data processing for the GPT-5.2 model t…
-
New method boosts AI model sensitivity to critical input edits
A new research paper introduces "abductive preference learning" (APL) to improve how vision and language models handle semantically critical input edits. Current models often ignore such edits, defaulting to their pre-t…
-
New benchmark aims to align LLM survey evaluators with human reviewers
Researchers have introduced SurveyReview, a new benchmark and dataset designed to evaluate large language models (LLMs) when they are used as survey evaluators. This benchmark addresses the lack of systematic alignment …