by Denkstrom
All storiesAI Market: Claude 4.8 Beats GPT-5.5 on Code

AI Market: Claude 4.8 Beats GPT-5.5 on Code

Anthropic has widened its lead on difficult coding tasks with Claude Opus 4.8: 69.2 percent on SWE-bench Pro, eleven percentage points ahead of GPT-5.5. Meanwhile, DeepSeek V4 Pro costs less than one-thirtieth of Claude after a 75 percent price cut.

Eighty-seven cents buys a million output tokens from Chinese AI model DeepSeek V4 Pro. Western competitors charge thirty times more. Yet Anthropic's Claude Opus 4.8 doesn't fall behind, thanks to a benchmark measuring real developer work where Claude beats OpenAI's GPT-5.5 by eleven percentage points. The AI market is fighting on two fronts: a quality competition among Western providers and a price war China has already won.

What Changed Since April

On April 24, 2026, Chinese startup DeepSeek released V4 Pro. With a context window of one million tokens and an initial price of 3.48 dollars per million output tokens, it was cheaper than Claude or GPT-5.5 from the start. Since May 31, DeepSeek has cut prices by 75 percent: one million tokens now costs 87 cents. That's a price level Western providers can't remotely match.

Five weeks after DeepSeek's launch, on May 28 Anthropic released Claude Opus 4.8. The model significantly improves on SWE-bench Pro. SWE-bench Pro uses real GitHub issues from open-source projects and forgoes prefabricated test cases: the model must recognize itself whether its solution works. Few benchmarks come closer to actual developer work. Claude 4.8 achieves 69.2 percent. GPT-5.5 reaches 58.6 percent, DeepSeek V4 Pro 55.4 percent.

Two Benchmarks, Two Stories

On the somewhat easier SWE-bench Verified, which uses existing test cases, the picture flips: DeepSeek V4 Pro 80.6 percent, Claude Opus 4.8 80.8 percent. The two models run virtually tied. A simple score comparison would suggest: DeepSeek and Claude are equal, but one costs one-thirtieth.

The difference between Verified and Pro explains the real difference. SWE-bench Verified tests known bugs with existing test harnesses. SWE-bench Pro tests blind: the model gets a problem, no hints, no checks. This difference forms the practical core. Anyone using AI for independent development tasks without human review needs Pro results. There Claude holds a 13.8 percentage-point edge over DeepSeek.

DeepSeek founder Liang Wenfeng proved since the R1 shock of January 2025 that Chinese AI is benchmark-capable. V4 Pro has open weights on Hugging Face and runs on Huawei chips rather than NVIDIA hardware according to the company. For developers outside the U.S. unwilling or unable to pay dollar prices, DeepSeek is no longer a second choice but a first.

Who Still Pays 25 Dollars Per Million Tokens

Despite the price gap, a stable market exists for expensive AI. A developer generating one million output tokens daily faces roughly 730 dollars monthly difference between DeepSeek and Claude. Yet many teams choose Claude or GPT-5.5.

First: performance gaps on hard tasks. The 13.8 point gap on SWE-bench Pro means Claude fails far less often on real engineering work. For simple tasks the difference is marginal. On autonomous agents working hours without human oversight, it compounds.

Second: compliance and geopolitics. U.S. firms in regulated industries (banking, pharma, defense) bet on vendors whose infrastructure falls under U.S. law. Security researcher Bruce Schneier noted in 2025 that running on Huawei chips says nothing about data protection in operation. European privacy regulators haven't definitively answered whether using DeepSeek APIs complies with GDPR.

Third: benchmark skepticism. Critics note SWE-bench Verified and Pro capture only part of real software development. Poolside AI, which builds developer tools on its own models, publicly stated it barely uses SWE-bench results because corporate codebases differ structurally from open-source repositories.

Anthropic's IPO Bet on Benchmark Leadership

Anthropic plans its IPO for October 2026. At a valuation around 965 billion dollars, the question of which model leads SWE-bench Pro is also an investor question. Who ranks first can justify higher enterprise prices and has direct sales advantage when Fortune 500 firms solicit AI for software development.

GPT-5.5 currently trails Claude 4.8 by 10.6 percentage points on SWE-bench Pro. OpenAI hasn't announced when a successor appears. DeepSeek must show whether its next model closes the 13.8-point Pro gap without abandoning the 87-cent price anchor. Until then, market division stays stable: those needing cheapest take DeepSeek. Those solving hardest problems take Claude.