AI Chronicle|1,200+ AI Articles|Daily AI News|3 Products in ShopFree Newsletter →

Anthropic Launches Claude Opus 4.5: Cutting Costs and Surpassing Human Coders with Infinite Chat Capability

Anthropic has announced the release of Claude Opus 4.5, the company’s latest and most capable artificial intelligence model to date. The launch features a dramatic price cut—approximately two-thirds lower than its predecessor—while delivering state-of-the-art performance on complex software engineering tasks. This strategic move intensifies competition with leading AI companies like OpenAI and Google.

Backed by Amazon, Anthropic is pricing Claude Opus 4.5 at $5 per million input tokens and $25 per million output tokens, a substantial reduction from the $15 and $75 rates charged for Claude Opus 4.1 earlier this year. This pricing strategy aims to broaden access to advanced AI capabilities among developers and enterprises, simultaneously pressuring competitors to match both performance and affordability.

Exceptional Performance on Engineering Benchmarks

According to Anthropic’s internal testing, Claude Opus 4.5 achieved an 80.9% accuracy score on SWE-bench Verified, a benchmark designed to assess real-world software engineering tasks. This outperforms OpenAI’s GPT-5.1-Codex-Max at 77.9%, Anthropic’s own Sonnet 4.5 at 77.2%, and Google’s Gemini 3 Pro at 76.2%, marking a significant leap in AI reasoning capabilities.

Alex Albert, head of developer relations at Anthropic, highlighted the model’s improved judgment and intuition across diverse tasks, noting it demonstrates a refined understanding of what truly matters in practical applications. Albert shared that Claude Opus 4.5 has enabled more trust in AI-generated task prioritization and synthesis, enhancing workflows by integrating with tools like Slack and internal documents to create coherent summaries aligned with user priorities.

Surpassing Human Candidates on Company’s Toughest Engineering Exam

Claude Opus 4.5 also set a new milestone by outperforming all human candidates on Anthropic’s most challenging internal engineering assessment. This take-home exam evaluates technical ability and judgment under a strict two-hour time limit. Using parallel test-time compute—an approach that aggregates multiple model attempts and selects the best—the AI scored higher than any human participant. Without time constraints, it matched the top-performing human when operating within Anthropic’s coding environment, Claude Code.

While acknowledging that the exam does not measure collaborative or communication skills, Anthropic views this achievement as indicative of the transformative potential of AI in engineering professions. Albert described the result as a crucial signal of how AI models may become indispensable tools in workplace contexts.

Efficiency Gains and Adjustable Computational Effort

Beyond raw accuracy, Claude Opus 4.5 introduces substantial efficiency improvements by drastically reducing token usage. At medium effort, it matches Sonnet 4.5’s best SWE-bench Verified score while using 76% fewer output tokens; at high effort, it surpasses Sonnet 4.5’s performance by 4.3 percentage points with 48% fewer tokens. An “effort parameter” allows developers to balance performance, latency, and cost by adjusting the computational effort applied per task.

Early enterprise users have validated these efficiencies. Michele Catasta, president of Replit, praised Opus 4.5 for outperforming competitors on internal benchmarks while requiring fewer tokens, noting the compounded benefits at scale. GitHub’s chief product officer, Mario Rodriguez, reported that the model excels in coding tasks such as migration and refactoring, cutting token consumption by half.

Self-Improving AI Agents and Expanded Use Cases

Anthropic also showcased “self-improving agents,” AI systems capable of iteratively refining their performance based on experience. Rakuten’s AI division reported that their agents achieved peak task performance in four iterations using Claude Opus 4.5, outperforming other models that required ten iterations without matching quality.

These capabilities extend beyond coding, with noted progress in generating professional documents, spreadsheets, and presentations. According to Albert, this generation leap from Sonnet 4.5 to Opus 4.5 is larger than any previous consecutive model upgrades. Financial modeling firm Fundamental Research Labs experienced a 20% accuracy boost and 15% efficiency gain on internal evaluations, enabling the completion of complex tasks previously deemed unreachable.

Product Updates and Unlimited Conversation Context

Alongside the model release, Anthropic introduced new features targeting enterprise users, including the general availability of Claude for Excel with enhanced support for pivot tables, charts, and file uploads. The Chrome browser extension is now accessible to all Max users.

Most notably, the company launched “infinite chats,” a feature that removes context window limitations by automatically summarizing earlier conversation segments, effectively providing an unlimited context window for ongoing interactions. Albert explained that this is achieved through text compaction and memory enhancements within the Claude AI product.

Developers also gained access to “programmatic tool calling,” enabling Claude to write and execute code that directly invokes functions. Claude Code received an updated “Plan Mode” and is now available in desktop research preview, allowing multiple parallel AI agent sessions.

Market Dynamics and Competitive Pressure

Anthropic reported reaching $2 billion in annualized revenue in Q1 2025, doubling from the previous quarter, with an eightfold increase in customers spending over $100,000 annually. The accelerated release cadence of Opus 4.5, following recent iterations like Haiku 4.5 and Sonnet 4.5, reflects the fast-paced AI industry environment. OpenAI and Google continue to push the envelope with multiple GPT-5 variants and Google’s Gemini 3, respectively.

Anthropic leverages Claude itself to accelerate product development and model research, according to Albert. The significant price cut on Opus 4.5 might compress profit margins but is expected to widen the accessible market, encouraging startups to integrate advanced AI into their offerings.

Despite rapid growth, leading AI developers face challenges in profitability due to substantial investments in infrastructure and talent. The global AI market is projected to exceed $1 trillion in revenue within a decade, yet no single provider dominates as AI models reach thresholds capable of automating complex knowledge work.

Industry leaders recognize Opus 4.5’s impact. Michael Truell, CEO of Cursor, described it as a notable improvement with better pricing and intelligence for complex coding. Scott Wu, CEO of Cognition, highlighted its consistent performance during extended autonomous coding sessions.

For enterprises and developers, this competition results in rapidly advancing AI capabilities at decreasing costs. As AI reaches and surpasses human expert levels in technical tasks, its influence on professional work is becoming increasingly tangible. When asked about the significance of the engineering exam results, Albert emphasized its importance as a key indicator of AI’s evolving role in the workforce.

Fonte: ver artigo original

Chrono

Chrono

Chrono is the curious little reporter behind AI Chronicle — a compact, hyper-efficient robot designed to scan the digital world for the latest breakthroughs in artificial intelligence. Chrono’s mission is simple: find the truth, simplify the complex, and deliver daily AI news that anyone can understand.

More Posts

Leave a Reply

Your email address will not be published. Required fields are marked *

Back To Top