On April 24, Chinese AI firm DeepSeek dropped V4, its long-awaited new flagship model, and honestly, it's a pretty big deal - even if it won't shake the AI world quite like its predecessor R1 did. V4 can process much longer prompts thanks to a nifty new design, it's open source (so anyone can download, use, and modify it), and it's making noise for reasons that go beyond just benchmarks.
First, the price. DeepSeek claims V4's performance rivals the best models at a fraction of the cost, which is great news for developers who don't want to mortgage their startups to pay for AI. V4 comes in two flavors: V4-Pro, a beefier model for coding and complex tasks, and V4-Flash, a smaller, speedier version. Both have reasoning modes that let the model show its work, like a math teacher who actually cares. Pricing? V4-Pro charges $1.74 per million input tokens and $3.48 per million output tokens - peanuts compared to OpenAI and Anthropic. V4-Flash is even cheaper at about $0.14 per million input tokens and $0.28 per million output tokens, making it one of the cheapest top-tier models around. That's the kind of price point that makes you want to build an app just because you can.
Performance-wise, V4 is a massive leap from R1, which is about as unsurprising as finding out water is wet. On major benchmarks, V4-Pro matches Anthropic's Claude-Opus-4.6, OpenAI's GPT-5.4, and Google's Gemini-3.1, and it beats open-source rivals like Alibaba's Qwen-3.5 and Z.ai's GLM-5.1 on coding, math, and STEM problems. It's also aces at agentic coding tasks and multistep problem-solving. DeepSeek's internal survey of 85 developers? More than 90% put V4-Pro in their top choices for coding. And for the agentic crowd, V4 is optimized for frameworks like Claude Code, OpenClaw, and CodeBuddy - because who doesn't want their AI to have a sidekick?
Now, the long-context thing. Both V4 versions can handle 1 million tokens - enough to fit all three volumes of The Lord of the Rings and The Hobbit combined. That's a lot of hobbits. But the real innovation is how DeepSeek did it: they tweaked the attention mechanism to be more selective, compressing older info and focusing on what matters now. This cuts computing power to 27% and memory to 10% compared to previous model V3.2 in a 1-million-token context (V4-Flash uses just 10% of compute and 7% of memory). That means you can build an AI coding assistant that reads your entire codebase without it forgetting what it read five minutes ago.
And now the geopolitics. V4 is DeepSeek's first model optimized for domestic Chinese chips, specifically Huawei's Ascend. This is a big deal because it's a test of whether China can wean itself off Nvidia. US export controls have cut off Chinese firms from Nvidia's best chips, so Beijing is pushing for a homegrown AI stack. DeepSeek reportedly gave early access only to Chinese chipmakers, and Huawei announced on Friday that its Ascend 950 supernodes will support V4. But let's not get ahead of ourselves: DeepSeek's technical report suggests they're using Chinese chips for inference, but training may still rely on Nvidia. Tsinghua professor Liu Zhiyuan told MIT Technology Review that only part of V4's training was adapted for Chinese chips, and multiple anonymous sources said Chinese chips are better for inference than training. Still, DeepSeek says V4-Pro prices could drop significantly once Huawei's Ascend 950 supernodes ship at scale in the second half of this year.
So, will V4 shake things up like R1? Probably not, but it's a sign that China is building a parallel AI infrastructure - and doing it with a price tag that'll make you do a double-take.
The Good Times
News in your inbox.
One sardonic roundup, delivered on your schedule. Free. Unsubscribe whenever your tolerance for wit runs out.
Already subscribed but we never reach your inbox? Check your spam folder and hit 'Not spam' (or 'Remove from spam') to bust us out of junk-mail purgatory. You'll be helping everyone else too.
Don't open any of our emails for a month and you'll be automatically removed from the mailing list.
Rewrite Article
Select parts to regenerate with a fresh AI pass. Translations will be updated automatically.
Generate AI Image
Creates a sardonic version of the article image using OpenAI.