GPT 5.4-Pro
PulseAugur coverage of GPT 5.4-Pro — every cluster mentioning GPT 5.4-Pro across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
AI models tackle unsolved math problems with formal proofs and new discoveries
Two new research papers explore the capabilities of large language models in advanced mathematical reasoning and discovery. The first paper introduces Magenta, a system that bridges informal natural language math proble…
-
MineBench 4.0 released with community gallery, iOS app, and private model testing
MineBench, a benchmark for evaluating AI models' ability to generate 3D structures, has released version 4.0. This update includes a community gallery for users to showcase and upvote custom prompts, and introduces A/B …
-
OpenCode Zen users face payment and geo-blocking issues
Users in Russia are facing difficulties accessing OpenCode Zen due to a combination of payment issues and geo-blocking. While the service advertises a pay-as-you-go model with no markups, users have reported that their …
-
GPT-5.4 Pro solves 60-year-old math problem, sparking debate
An amateur mathematician utilized GPT-5.4 Pro to solve a long-standing mathematical problem posed by Paul Erdős, which had eluded human mathematicians for 60 years. The AI's approach differed significantly from human me…
-
New VLM framework boosts 3D view planning with self-exploration
Researchers have developed a new framework to improve the view planning capabilities of Vision-Language Models (VLMs) in 3D environments. The proposed method alternates self-exploration with view graph distillation, whe…
-
LLMs Overwhelmingly Reproduce Majority Human Grading in Thai Bar Exam Study
A new study on the Thai bar examination reveals that while human examiners sometimes diverge on grading free-form essays due to ambiguous rubric interpretations, Large Language Models (LLMs) overwhelmingly converge on t…
-
Tiny models outperform frontier AI in agent coding benchmark
A recent agent coding benchmark revealed that smaller, more efficient models are outperforming larger, frontier models. The SmolLM3 3B model, capable of running on a laptop, achieved a score of 93.3, significantly surpa…
-
Google DeepMind AI assists mathematicians, tops FrontierMath benchmark
Google DeepMind has released an AI system called "AI Co-Mathematician" designed to collaborate with human mathematicians on complex problems. This system, built on Gemini 3.1 Pro, achieved a new state-of-the-art score o…
-
GPT-5.4 Pro assists 23-year-old in solving 60-year-old Erdős problem
A 23-year-old individual leveraged GPT-5.4 Pro to solve a 60-year-old mathematical problem known as an Erdős problem. The solution, a proof based on discrete Markov chains, has been verified using the Lean proof assista…
-
GPT-5.5 Pro excels on benchmarks; Microsoft Playwright aids web agents
OpenAI's GPT-5.5 Pro has reportedly achieved significant gains on the Epoch benchmark, with its base version outperforming the previous Pro model. This suggests substantial efficiency improvements in OpenAI's latest ite…
-
OpenAI's GPT-5.4 Pro assists in solving a 60-year-old mathematical problem
OpenAI has announced that a 60-year-old unsolved mathematical problem, known as the Erdős problem, has been solved with the assistance of their GPT-5.4 Pro model. Researchers from OpenAI detailed how the AI's advanced r…
-
OpenAI's GPT-5.4 Pro aids in solving 60-year-old math problem
OpenAI's GPT-5.4 Pro assisted in solving a 60-year-old mathematical problem posed by Paul Erdős. This development raises questions about the future of AI in mathematical research and problem-solving. OpenAI researchers …
-
Amateur uses ChatGPT to solve 60-year-old math problem, surprising experts
A 23-year-old amateur mathematician named Liam Price has solved a 60-year-old mathematical problem, known as an Erdős problem, using ChatGPT. Price, who has no advanced mathematics training, reportedly used a single pro…
-
AI and human collaboration fully resolves Donald Knuth's "Claude Cycles" problem
An AI startup has collaborated with Professor Donald Knuth on his "Claude Cycles" problem, a graph decomposition conjecture. Initially, an AI system named Claude Opus 4.6 explored the problem for about an hour, leading …