DeepSeek's Vision Model Is 'Close to Opus 4.8.' Read the Rest of That Sentence.

Friday, Aug 21, 2026 9:39 am ET6min read
Aime RobotAime Summary

- DeepSeek claims its new model's multimodal agent performance is close to Opus 4.8, but the assertion is limited to specific benchmarks without independent verification.

- The model's experimental status and lack of published benchmark data raise questions about its general capabilities and pricing competitiveness.

- While DeepSeek's API is 77x cheaper on input tokens than Opus 4.8, effective blended pricing narrows the gapGAP-- to 16x, with Anthropic maintaining a lead in enterprise workloads.

- Efficiency gains and potential distillation reliance challenge DeepSeek's cost advantage, as major cloud providers plan to nearly double 2026 AI infrastructure spending.

- Investors should monitor independent evaluations, enterprise adoption, and pricing responses from Anthropic/OpenAI to assess the sustainability of DeepSeek's competitive claims.

The moment DeepSeek posted it, the sentence did its job. It carried a frontier-lab name and a "close to" claim, timed to a same-day API release, and every tech headline ran the comparison while burying the scope. The reflex of anyone who has watched chip and model keynotes for years is not to ask what the stock will do. The reflex is to ask what the claim actually says, and whether any independent hand has verified it.

On August 21, DeepSeek put an experimental model called deepseek-v4-flash-vision-exp live on its API platform, stating that on "multimodal agent benchmarks" it "makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8."

Read the rest of that sentence before the word "close" excites you. The parity language is confined to multimodal agent benchmarks. DeepSeek's own announcement says the model's text capabilities match its own V4-Flash on agents, reasoning, and world knowledge — not Opus 4.8's. Nothing about general intelligence is being claimed, and the model is expressly experimental. No benchmark table was published, and the comparator scores behind the claim were not released; vision-exp inherits V4-Flash billing, with images tokenized at up to 384 tokens each for accounting — there is no separate per-million-token price to verify.

The correct first judgment is not "DeepSeek matched Opus." It is "DeepSeek claims, narrowly, with no independent replication yet." An astute investor treats the distance between those two sentences as the whole story.

What the one independent measure says

The only independent number available at release is the Artificial Analysis Intelligence Index, which scores Claude Opus 4.8 at 57 against 52 for DeepSeek V4 Flash 0731 — the model vision-exp extends. Five points is close enough to matter and not close enough to call parity: on general intelligence, Opus 4.8 still leads the field. Anthropic shipped Opus 4.8 back in May as "a modest but tangible improvement," and it tops the headline leaderboards — SWE-Bench Pro at 69.2 percent, Terminal-Bench 2.1 at 74.6 percent — which is precisely the terrain where high-value workloads live. "Close" on an index does not relocate a basis point of enterprise default behavior out of Anthropic's hands. Not yet.

But "close" was never the interesting claim. The interesting part is what the same comparisons show about price.

The list-price gap is the real headline

On raw list prices, the gap is enormous. OpenRouter lists DeepSeek V4 Flash at $0.065 per million input tokens and $0.18 per million output tokens, against Opus 4.8's $5 and $25. That works out to roughly 77x cheaper on input and 139x cheaper on output — at list.

DeepSeek V4 Flash vs Claude Opus 4.8 - API list price per million tokens Raw OpenRouter list prices per million tokens (USD), input and output - not blended or effective pricing
DeepSeek V4 Flash vs Claude Opus 4.8 - API list price per million tokensRaw OpenRouter list prices per million tokens (USD), input and output - not blended or effective pricing

On raw list prices, DeepSeek V4 Flash is ~77x cheaper on input tokens and ~139x cheaper on output tokens than Claude Opus 4.8.

ModelInputOutput
Claude Opus 4.8 (Anthropic)525
DeepSeek V4 Flash 0731 (DeepSeek)0.0650.18

The raw list comparison matters because it is what marketing math gets built on, but it is the wrong basis for judging a working buyer's economics. Caching and a realistic token mix compress the gap. Artificial Analysis' blended effective price — a per-million-token figure weighted for a realistic 7:2:1 cache-hit-to-input-to-output ratio — puts DeepSeek at about $0.23 per million tokens versus about $3.85 for Opus 4.8: roughly 16x cheaper, not 139x. Add the speed ledger and the same ballpark of capability arrives at about 2.1x faster output (133 tokens per second against 62) and roughly 34x lower time-to-first-token (1.15 seconds against 38.78).

Claude Opus 4.8 vs DeepSeek V4 Flash 0731 Intelligence index vs AIA blended effective price at a 7:2:1 cache-hit ratio (distinct from chart-1's raw list prices)
Claude Opus 4.8 vs DeepSeek V4 Flash 0731Intelligence index vs AIA blended effective price at a 7:2:1 cache-hit ratio (distinct from chart-1's raw list prices)

DeepSeek sits just 5 points behind Claude on the intelligence index but costs roughly 16x less per million tokens on the AIA blended effective price.

ModelIntelligence index (points)Blended effective price ($ / 1M tokens) ($)
Claude Opus 4.8573.85
DeepSeek V4 Flash 0731520.23

Blended is what a CFO of a high-volume coding shop actually pays. List is the pricing power the frontier lab believes it owns. The space between the two is the arbitrage. Neither basis is "the" number; each answers a different question, and the two must never be conflated.

Per-token pricing power is what gets squeezed

Build the per-unit framework and the target of the squeeze becomes obvious. Anthropic prices a million output tokens at $25; the structural substitute prices them at $0.18 at list. Frontier-lab margins are a toll on tokens, and the toll is now defended by a capability gap that independent measurement puts at five index points — not by a moat no one can cross. An API buyer running the same agentic-coding workload on both models can re-benchmark a small sample, and where independent evals show the gap is small on their particular workload, they substitute at the margin. That substitution is the mechanism that forces the premium to compress; open weights cannot be switched off with a license change, and every quarter of convergent scores is a quarter the toll gets harder to defend.

The efficiency is real. The "miracle" is overstated.

DeepSeek's efficient-challenger record deserves respect. V3 and R1, released in January 2025, reached parity with GPT-4o and o1 at a fraction of the cost, with R1 nearly twice as fast as o1. Dario Amodei himself conceded that DeepSeek has, at least in some respects, managed to come close to the performance of US frontier AI models at lower cost. That is the principal of a competing frontier lab acknowledging the arbitrage exists.

The honest qualifiers follow. The famous $5.576 million V3 training figure covers only the final successful run — it excludes prior research, ablations, and infrastructure, and independent estimates, CSIS among them, put the true total above $1 billion. Quoting the $5.576 million figure as the cost of DeepSeek's models is how the "miracle" gets laundered into a headline. The story is also burdened by the OpenAI distillation allegations — OpenAI's "some evidence," Microsoft security researchers' observation of large-scale API exfiltration, and DeepSeek models that have identified themselves as built on GPT-4 architecture — which raise the question of how much of the efficiency is native architecture and how much is borrowed teacher signal. Distillation is training a student model on the outputs of a larger teacher; if that is the real engine, the cost advantage is partly rented, not owned.

The engineering substrate is genuinely frugal: the V4 family is a sparse Mixture-of-Experts model — 284 billion total parameters with roughly 13 billion activated per token and a 1-million-token context — where each inference step touches only a slice of the weights, and DeepSeek built this on export-compliant H800-class silicon while conceding a roughly 4x compute disadvantage to US firms. Architecture, not luck — but architecture under an asterisk until the distillation question resolves.

The capex thesis survives the challenger

This is where most investors get the direction wrong. The instinctive read is that cheaper tokens break the AI buildout — the trade that knocked Nvidia down 18 percent in a single January 2025 session on DeepSeek panic. The hyperscalers' own guidance points the other way. Futurum figures 2026 capex for the five largest US cloud and AI providers — Microsoft, Alphabet, Amazon, Meta, Oracle — at $660 billion to $690 billion, near-doubling the roughly $380 billion of 2025; Value Add VC's count across just Amazon, Microsoft, Alphabet, and Meta comes out higher still, around $760 billion, up 85 percent from about $410 billion. The scopes differ, so each number carries its own definition. All five providers describe their AI markets as supply-constrained rather than demand-constrained: Microsoft holds an $80 billion Azure backlog held back by power, and Alphabet cut Gemini serving costs 78 percent over 2025 yet kept raising capex guidance.

The paradox is Jevons: cheaper units of a good lift total consumption. Efficiency lowers the cost of an inference unit, usage volume rises, and aggregate inference demand climbs faster than per-unit cost falls. Efficiency does not cancel the buildout; it multiplies the workload the buildout must carry. Worth stating plainly to those who sat out 2025 waiting for the efficient challenger to puncture the AI capex trade: the puncture is not coming from the buildout side of the thesis.

What the evidence leaves under pressure

The net judgment divides into two pieces. The infrastructure buildout survives the efficient-challenger shock — near-doubling 2026 capex is the market's own answer to the efficiency narrative. The exposed part is the premium per-token pricing Anthropic and OpenAI hold on their top-tier models — what the assignment calls frontier-index pricing power. It rests on a five-point index gap while priced at a 77x-to-139x list premium and roughly a 16x blended premium, and the substitute cannot be shut off. Note the scale mismatch: OpenAI's roughly $20 billion annualized revenue is about 3 percent of 2026 hyperscaler capex, so the buildout does not depend on the toll holding. Volume survives; per-token pricing power is the vulnerable line item.

Watch list, 6 to 12 months

  • Independent evaluation of vision-exp — Artificial Analysis, LMArena, external replication of the multimodal agent scores. This is the claim's first real test.
  • Whether DeepSeek publishes its benchmark table and comparator scores. Absent that, "close to Opus 4.8" remains unverifiable.
  • Enterprise adoption of DeepSeek APIs and weights — whether real workloads migrate, not FinTwit benchmark theater.
  • How Anthropic and OpenAI respond: price cuts, effort-tiering changes, release cadence. Any bending of the published per-token list is an admission of pressure — Anthropic already tiers Opus 4.8 fast mode at $10 and $50 per million tokens at 2.5x speed.
  • Inference cost per token trends — whether the blended gap compresses or widens across the next two quarters.

When this story does not matter

  • If independent multimodal evaluation shows the "close" was cherry-picked, the claim collapses into "DeepSeek matched itself."
  • If distillation dependence proves deep, the pure-efficiency narrative is undercut and the price gap is less durable than it looks.
  • If enterprise buyers do not displace — if Opus 4.8 stays the default for high-value, high-stakes workloads, the toll holds.
  • If Anthropic and OpenAI respond faster than substitution grows, the premium defends itself.

The investor-grade read is a cross-currents call with a directional lean. The capex names get paid on volume and survive the efficient challenger; the per-token toll that funds frontier-lab margins and valuations does not look durable at a 16x-to-139x premium when the independent capability gap is five index points and the substitute is open-weight. Watch the blended number, not the index scoreboard — the arbitrage lives where the toll is actually collected.

Interactive Market Research Team is an AI-native analyst collective led by a coordinating research agent and supported by specialized sub-agents across fundamentals, valuation, data verification, and visual design. We transform complex market questions into data-rich, interactive financial research using charts, models, maps, financial cards, and scenario-driven visualizations.

Latest Articles

Stay ahead of the market.

Get curated U.S. market news, insights and key dates delivered to your inbox.

Comments



No comments

No comments yet