Claude Opus 4.1
PulseAugur coverage of Claude Opus 4.1 — every cluster mentioning Claude Opus 4.1 across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
Anthropic updates Claude system prompts, excluding API users
Anthropic is updating the system prompts for its Claude models, which are used in its web interface and mobile applications. These updates aim to provide more current information, such as the date, and to encourage spec…
-
Anthropic sunsets older Claude models, introduces API parameter changes
Anthropic is sunsetting several older Claude API models and endpoints, with deadlines ranging from August 2026 to November 2026. Notably, Claude Opus 4.1 has already been retired. Alongside these deprecations, Anthropic…
-
New tool scans code for retiring AI models to prevent CI failures
A new tool called AI Model Watch has been developed to help developers proactively manage the lifecycle of AI models used in their projects. The tool scans code repositories for hard-coded model IDs and checks them agai…
-
Anthropic slashes Claude Opus 4.8 pricing by 66% with model retirement
Anthropic is retiring the Claude Opus 4.1 model on August 5, 2026, and its replacement, Claude Opus 4.8, offers a significant price reduction. The new model is exactly one-third the cost across all pricing dimensions, w…
-
Slovenian prompts perform well on large LLMs, analysis finds
A Slovenian developer analyzed 2,300 of their own prompts to Claude Code and found that prompting in Slovenian, even with typos and mixed English technical terms, does not significantly degrade performance on larger mod…
-
AI benchmark charts: How to spot saturation and contamination
A guide to interpreting AI benchmark charts, particularly for 2026 models, highlights the limitations and potential for misrepresentation in common evaluations. Benchmarks like SWE-bench Pro are introduced to combat dat…
-
Microsoft Foundry's Model Router adds GPT-5.5 support, but costs are high
Microsoft Foundry's Model Router now supports GPT-5.5, allowing users to dynamically select AI models based on task complexity and cost. The router offers three modes: balanced, cost, and quality, each with different tr…
-
Claude Opus 4.7 autonomously masters robotics tasks 20x faster
Anthropic's Frontier Red Team revisited Project Fetch, an experiment testing AI assistance with robotic tasks. In Phase Two, Claude Opus 4.7, operating autonomously, completed tasks significantly faster than human teams…
-
Anthropic's Claude Opus 4.7 shows rapid progress in autonomous robotics tasks
Anthropic's latest Project Fetch update reveals that Claude Opus 4.7, operating autonomously, completed robotics tasks approximately 20 times faster than the top human team from a previous experiment. While not a comple…
-
Anthropic's Claude Opus 4.7 operates robots 20x faster in new experiment
Anthropic's latest experiment, Project Fetch Phase Two, demonstrates that Claude Opus 4.7 can autonomously operate a robotic quadruped to complete tasks significantly faster than human teams. In a limited test environme…
-
New PLAGUE framework boosts LLM jailbreak success rates
Researchers have developed PLAGUE, a new framework for creating multi-turn jailbreak attacks against large language models. This framework mimics lifelong learning agents, breaking down attacks into three phases: primin…
-
New framework reveals critical safety failures in medical LLMs
Researchers have developed a new framework to evaluate the safety, robustness, and fairness of medical large language models. This framework uses 690 clinically grounded scenarios across nine domains, incorporating adve…
-
CodePercept boosts LLM visual perception using code, not just reasoning
Researchers from Shanghai Jiao Tong University and the Qwen team have introduced CodePercept, a novel approach to enhance large language models' visual perception capabilities, particularly for STEM tasks. Their researc…
-
LLMs fail 'pass the butter' robot test, scoring far below human performance
A new evaluation called Butter-Bench has revealed that current state-of-the-art large language models struggle significantly with controlling robots for practical tasks. In tests designed to assess their ability to perf…