PulseAugur
EN
LIVE 16:50:37

User finds Claude Opus 5 unreliable for planning tasks despite initial success

A user initially found Anthropic's Claude Opus 5 to be frustratingly verbose and inefficient, similar to other users' experiences. However, during a project involving CAN networking, Opus 5 unexpectedly performed well, successfully identifying errors made by Sonnet and proposing effective solutions. This positive experience was short-lived, as a subsequent attempt to create a simple skill using Opus 5 resulted in a disastrously complex and token-intensive process involving multiple agents, leading the user to question Opus 5's reliability for planning tasks. AI

IMPACT Highlights potential inconsistencies in large language model performance and user experience.

RANK_REASON User experience report on a specific model's performance.

Read on r/Anthropic →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

User finds Claude Opus 5 unreliable for planning tasks despite initial success

COVERAGE [1]

  1. r/Anthropic TIER_1 English(EN) · /u/Vibroverbus ·

    "Hey look Opus 5 is not so bad... NO WAIT WTF IT IS HORRIBLE...."

    <!-- SC_OFF --><div class="md"><p>I've struggled with Opus 5 on/off like others. I get the frustrating crazy verbose crazy inefficient token black hole behavior, and then flip to 4.6 sometimes but then end up trying 5 again with different prompting style or assignment in hopes I …