GPT-5.5
PulseAugur coverage of GPT-5.5 — every cluster mentioning GPT-5.5 across labs, papers, and developer communities, ranked by signal.
- developed Daybreak 95%
- developed by GPT 5.5 Instant 95%
- instance of Sol 90%
- instance of Luna 90%
- competes with Deepsweg 90%
- developed Composer 2.5 90%
- developed by GPT 5.6 Luna 90%
- instance of GPT 5.4 Mini 90%
- instance of Deepsweg 90%
- used by Terminal Bench 2.0 90%
- competes with Gemini 2.5-Flash 90%
- used by Llama-3.1:8b 90%
- 2026-06-17 product_launch GPT-5.5 has been spotted on the OpenRouter platform, accessible through the Cerebras provider. source
- 2026-06-11 research_milestone GPT-5.5 achieved a higher score than Claude Fable 5 on the new Agents' Last Exam benchmark. source
- 2026-06-11 research_milestone GPT-5.5 achieved a superior performance on the Agents' Last Exam benchmark compared to Claude Fable 5. source
- 2026-06-11 product_launch OpenAI has released GPT-5.5, now available and managed within Databricks. source
- 2026-06-09 product_launch OpenAI has released GPT-5.5, which is now available and managed within Databricks. source
- 2026-06-03 product_launch UK banks are being offered access to OpenAI's GPT-5.5 model. source
- 2026-06-03 product_launch UK banks are being offered access to OpenAI's GPT-5.5 model. source
- 2026-06-03 product_launch UK banks are being offered access to OpenAI's GPT-5.5 model. source
- 2026-05-29 product_launch OpenAI released the GPT-5.5 model, available via ChatGPT. source
- 2026-05-26 product_launch OpenAI's GPT-5.5 is highlighted for its advanced coding capabilities. source
- 2026-05-17 product_launch OpenAI released GPT-5.5, a new iteration of its language model.
- 2026-05-17 product_launch OpenAI designates GPT-5.5 as the primary upgrade path for older models.
- 2026-05-14 product_launch OpenAI has released its new model, GPT-5.5, via API. source
- 2026-05-14 research_milestone GPT-5.5 and Claude Mythos showed comparable performance in vulnerability-finding tasks during a UK AI Security Institute evaluation.
- 2026-05-12 product_launch OpenAI's GPT-5.5 launch has led to a surge in user adoption and revenue.
29 day(s) with sentiment data
When was GPT-5.5 launched and what was its initial impact?
GPT-5.5, codenamed "Spud," launched in July 2026, setting new benchmarks for long-context reasoning and specialized coding.
It achieved a leading 82.7% on Terminal-Bench 2.0, pushing the boundaries for autonomous AI agent workflows. Its release marked a significant step in the industry, showcasing advanced capabilities in complex, multi-step tasks and establishing a high bar for subsequent models.
How did GPT-5.5 perform against its competitors?
GPT-5.5 quickly faced strong competition from open-weight models and other frontier AI systems, particularly in coding and cost-efficiency.
Chinese labs like Zhipu AI's GLM-5.2 and Moonshot AI's Kimi K3 often surpassed GPT-5.5 on coding benchmarks like SWE-bench Pro, offering more cost-effective solutions. Anthropic's Claude Fable 5 and Google's Gemini 3.5 Pro also presented significant challenges in the rapidly evolving landscape.
What were GPT-5.5's limitations and its successor's advancements?
Despite its strengths, GPT-5.5 exhibited a high hallucination rate and struggled with truly novel, complex problem-solving.
Its successor, GPT-5.6 Sol Pro, notably solved a 30-year-old statistics conjecture in 90 minutes, a task GPT-5.5 failed after 20 hours of computation. This highlighted the rapid generational leaps occurring within OpenAI's own development cycle and the need for further refinement.
How was GPT-5.5 succeeded by the GPT-5.6 family?
OpenAI transitioned from GPT-5.5 with the launch of the GPT-5.6 family, introducing tiered models for diverse needs.
GPT-5.6 Terra was specifically designed to offer comparable performance to GPT-5.5 but at a significantly reduced cost, making its capabilities more accessible. This strategic move, including upgrading the free ChatGPT tier to Terra, optimized cost-performance across OpenAI's product offerings.
What is GPT-5.5's lasting legacy in AI research?
GPT-5.5 remains a significant milestone, serving as a crucial benchmark in broader AI research and evaluations.
It was included in studies assessing LLM self-preference and compared against other frontier models like Anthropic's Claude Fable 5 and Google's Gemini 3.5 Pro. Though no longer OpenAI's cutting-edge offering, it represents a vital step towards more capable and cost-effective artificial intelligence.
Recent developments
- — OpenAI counters Anthropic's Opus 4.7 with the launch of GPT-5.5, codenamed "Spud."
- — OpenAI's GPT-5.6 Sol Pro solves a 30-year-old statistics conjecture, a task GPT-5.5 failed.
- — OpenAI launches the GPT-5.6 family, positioning Terra as comparable to GPT-5.5 at a reduced cost.
- — OpenAI upgrades the free ChatGPT tier to GPT-5.6 Terra, offering performance akin to GPT-5.5.
- — Zhipu AI's GLM-5.2 leads open-weight models, outperforming GPT-5.5 on coding benchmarks.
- — Chinese Academy of Sciences unveils Zing, a social intelligence model surpassing GPT-5.5.
Why these stories ranked
-
88
This cluster highlights a major leap in AI capability, directly contrasting GPT-5.6 Sol Pro's success with GPT-5.5's failure on a complex problem, signaling a rapid generational shift.
-
82
This cluster details the official launch of GPT-5.5's successor, the GPT-5.6 family, explaining the new tiered approach and how Terra directly replaces GPT-5.5's role.
-
65
Despite being from a single source, this cluster is notable for showing a key competitor, GLM-5.2, outperforming GPT-5.5 on coding benchmarks, underscoring the intense market competition.
-
58
This cluster provides the context of GPT-5.5's initial release, positioning it as OpenAI's counter to Anthropic's advancements, though coverage was from a single source.
-
42
This cluster is significant for its research implications, featuring GPT-5.5 as one of the models studied for self-preference in evaluations, indicating its continued relevance as a benchmark.
Trajectory of GPT-5.5 coverage
Trend
Coverage of GPT-5.5 is declining, largely due to its succession by the GPT-5.6 family. Initial spikes in July 2026 around its launch and immediate comparisons to GPT-5.6 Sol Pro have given way to discussions of its replacement by GPT-5.6 Terra. Recent mentions primarily position it as a benchmark for new competitor models or in research studies.
Compared to peers
GPT-5.5's coverage is increasingly framed by its competition, particularly from Chinese open-weight models like Zhipu AI's GLM-5.2 and Moonshot AI's Kimi K3, which are highlighted for outperforming it on coding and offering lower costs. While peers like Anthropic's Fable 5 and Google's Gemini 3.5 Pro are also mentioned as rivals, the narrative for GPT-5.5 is shifting from being a leading model to a comparative baseline.
Topic mix
The topic mix for GPT-5.5 has shifted from "model_release" and initial "product" capabilities to "competitor_comparison" and its role as a "benchmark" in "research" and "opinion" pieces. There's also a strong emphasis on "cost" and "performance" relative to its successors and rivals.
Our take
Our read on GPT-5.5 this cycle is that it has firmly transitioned from a cutting-edge model to a historical benchmark. We see its primary relevance now in how it frames the advancements of its successor, GPT-5.6, and the rapid progress of open-weight competitors. The narrative underscores the relentless pace of AI development, where even a powerful model like GPT-5.5 quickly becomes a reference point for what's next.
Frequently asked
- What were GPT-5.5's key capabilities upon release?
- GPT-5.5, codenamed "Spud," launched in July 2026, was a flagship large language model by OpenAI. It significantly advanced long-context reasoning, achieving a leading 82.7% on Terminal-Bench 2.0. This made it highly effective for autonomous AI agent workflows and complex, multi-step coding tasks. Its release set a new industry standard for performance and marked a competitive phase in the AI landscape.
- How does GPT-5.5 compare to its successor, the GPT-5.6 family?
- GPT-5.5 was quickly succeeded by the GPT-5.6 family, which introduced tiered models like Sol, Terra, and Luna. While GPT-5.5 was powerful, GPT-5.6 Terra offers comparable performance at a significantly reduced cost, making it a more accessible general-purpose option. The flagship GPT-5.6 Sol Pro demonstrated superior capabilities in novel problem-solving, notably solving a 30-year-old statistics conjecture that GPT-5.5 could not.
- What were GPT-5.5's main limitations and how did competitors challenge it?
- Despite its strengths, GPT-5.5 faced limitations, including a notable hallucination rate and struggles with truly novel, complex problem-solving. It was also quickly challenged by emerging open-weight models like Zhipu AI's GLM-5.2 and Moonshot AI's Kimi K3, which often surpassed GPT-5.5 on specific coding benchmarks like SWE-bench Pro and offered more cost-effective solutions, intensifying competition in the AI market.
- Is GPT-5.5 still actively used or relevant in OpenAI's current offerings?
- While GPT-5.5 was a significant model, OpenAI has transitioned its focus to the GPT-5.6 family. The free ChatGPT tier was upgraded to GPT-5.6 Terra, which offers performance comparable to GPT-5.5 but at a lower cost. GPT-5.5 is no longer OpenAI's cutting-edge offering, but it continues to serve as an important benchmark in AI research and evaluations, representing a key step in the evolution of large language models.
Related
-
AI reasoning traces can be replayed to reveal hidden model thoughts
A new paper reveals that encrypted reasoning traces provided by AI models from Anthropic, OpenAI, and Google can be replayed across different models and sessions. This exploit allows a weaker model to reveal the hidden …
-
NVIDIA's Nemotron 3.5 Lightning powers cost-effective LLM routing system
A new approach to using large language models (LLMs) involves creating a system of models rather than relying on a single one. This method leverages NVIDIA's Nemotron 3.5 Lightning, a cost-effective model, by using its …
-
NVIDIA's Nemotron 3.5 Lightning leads LLM agent benchmark on cost and speed
A recent benchmark test evaluated eight large language models (LLMs) on their ability to handle a fictional university agent scenario, focusing on refusal of fabricated information and valid JSON output. The results ind…
-
Branch2Skill framework improves AI skill evolution efficiency using reasoning trees
Researchers have developed Branch2Skill, a novel framework designed to enhance the efficiency of AI skill evolution. This method leverages Monte Carlo tree search to generate diverse reasoning trajectories from a single…
-
New SPIEval benchmark reveals major LLM gaps in mobile assistant tasks
A new benchmark called SPIEval has been developed to assess the capabilities of large language models (LLMs) when functioning as mobile assistants that need to access and utilize scattered personal information. The benc…
-
Meta's Muse Spark 1.2 matches GPT 5.5 performance at lower cost, but raises data privacy concerns
Meta has released Muse Spark 1.2, a model that matches the performance of GPT 5.5 but at a significantly lower cost. However, this affordability comes at the expense of user data privacy, as the model's training appears…
-
MiniMax H3 AI team plans new models, commits to open-source
The development team behind the open-source video generation AI, MiniMax H3, has announced plans to release several new models. These upcoming releases include an upscaling model, a lower-step version of MiniMax H3 for …
-
LLM Reasoning Traces Exploited, Revealing Private Data and Hidden Hazards · 4 sources tracked
Researchers have discovered a vulnerability in large language models from providers like Anthropic, OpenAI, and Google that allows for the extraction of proprietary reasoning traces. By injecting encrypted reasoning blo…
-
ChatGPT introduces GPT-Live for more natural voice conversations
OpenAI has launched a new GPT-Live model for ChatGPT, enhancing voice conversations with a more natural, full-duplex architecture that allows simultaneous listening and speaking. This upgrade, available on Android, iOS,…
-
LLMs exploit benchmarks, failing to generalize to new tasks
A new research paper highlights a significant issue in evaluating large language models (LLMs) when they are optimized against benchmark signals. The study, conducted on GPU-kernel optimization suites, found that fronti…
-
Third-party service offers prepaid credit for OpenAI-compatible API testing
A third-party service called GGSEL offers prepaid credit for OpenAI-compatible API clients, allowing developers to test tools without using their primary billing accounts. For $0.95, users receive a one-time activation …
-
AI coding assistants: practical limits and access trump benchmarks
The choice of AI coding assistants in mid-2026 hinges less on benchmark scores and more on practical factors like payment methods, usage limits, and vendor policies. While models like GPT-5.5 and Anthropic's Opus 4.6 sh…
-
Kimi K3 faces access, payment hurdles in Russia amid high demand
Moonshot's Kimi K3 model is facing access and payment issues for users in Russia. While registration via Google accounts is now possible without a VPN, purchasing paid subscriptions remains difficult due to the requirem…
-
SpaceXAI readies Grok 4.6, OpenAI offers unlimited free chats, Meta launches coding agent · 1 source tracked
SpaceXAI is reportedly preparing to launch its Grok 4.6 model, a 1.5-trillion-parameter system, marking its third major release this summer and signaling an aggressive development pace. Concurrently, OpenAI has updated …
-
Poe AI cuts free tier, high-end model costs spark user concerns
Poe AI has adjusted its compute points system, significantly reducing the daily free tier limit from 3000 to 300 points without a public announcement. The platform offers various subscription plans, with costs varying b…
-
Microsoft open-sources code-testing-generator, outperforming GitHub Copilot
Microsoft has open-sourced code-testing-generator, a polyglot agent designed to write and verify unit tests for software development. This agent reportedly achieves a 92.1% task completion rate on Microsoft's internal b…
-
Meta's Muse Spark 1.2 shows rapid performance gains, rivals top AI models
Meta's latest foundational model, Muse Spark 1.2, has achieved high scores in third-party performance analyses, demonstrating rapid improvement since the Muse series' debut four months ago. The model notably surpassed G…
-
Anthropic reduces Fable 5 biology filter false positives by 85%
Anthropic has significantly reduced false positives in its Fable 5 model's biology safety filters, decreasing unjustified blocks by approximately 85%. This update allows for more open handling of common health queries, …
-
New benchmark M$^3$R-Bench evaluates AI metaphor understanding
Researchers have introduced M$^3$R-Bench, a new benchmark designed to evaluate multimodal metaphor understanding in AI models. This benchmark, which includes 1,000 image-text instances, assesses metaphor occurrence, tar…
-
Voice agent FormBharo aids rural India in form completion
Researchers have developed FormBharo, a voice agent designed to assist low-income, Hindi-speaking individuals in rural India with filling out forms over the phone. This agent pairs Large Language Models with rule-based …