Part 1: DeepSeek's V4 makes Chinese AI labs look like one mega-lab
The price implications are near-term, long-term is shared R&D logic.
Apologies for the delay on this update, as I deliberately wanted to wait for the bullets to fly for a bit, “让子弹飞一会.”
I have read through at least twenty reports on pricing and performance, but what dawned on me when I took a few days away from it all and thought through the ecosystem and disruption, is that V4’s implication on the business and overall sector that is very much overlooked right now. Thus, on a Saturday morning, the light bulb finally went off for me, and I jotted everything down. Part 1 will focus on V4’s direct implications for the industry, the signal behind the action, and how China’s open-source strategy is becoming increasingly cohesive. Part 2 looks at what V4 means for the tech stack and what Jesnen Huang keeps warning about across what he calls the five-layered cake.
The easiest way to understand Chinese AI right now is that the labs are, in some strange way, starting to work like different departments inside one large AI company.
Obviously, this is not literally true. DeepSeek, Z.ai, Moonshot, MiniMax, Alibaba, Tencent, and ByteDance still compete for users, enterprise customers, talent, capital, and attention. Fundraising is competitive, API pricing is competitive, and model launches are competitive. But technically, the ecosystem is becoming more collaborative than people realize. And it’s looking more like what we heard from the researchers in February, where DeepSeek now plays an infrastructure provider role in some sense.
DeepSeek is increasingly doing the foundation-layer work: architecture, efficiency, inference optimization, and now Huawei-stack adaptation. Z.ai is leaning into coding and enterprise deployment. Moonshot Kimi is leaning into long context and agentic orchestration. MiniMax is leaning into multimodality and low-cost inference. Alibaba, Tencent, and ByteDance are leaning into distribution, cloud, workflow, and product integration. So even though these companies are still fighting commercially, the technical layer is starting to look more modular. Each lab is leaning into its own area of expertise, while absorbing what others release as open weights, disclose through technical reports, or expose through benchmarks and deployment behavior.
This is where DeepSeek V4's implication is being overlooked. The obvious read is that V4 is another pricing shock, and that is not wrong, but obviously not as big a shock to the global AI community.
The second read is that it will pressure Z.ai, MiniMax, Kimi, and everyone else trying to charge premium prices for text, coding, and agentic workloads. On top of already very competitive listing price for its API, DeepSeek is currently offering a 75% discount on V4-Pro to developers first until May 5, and then extended to the end of May, and cut prices for input-cache hits across its API lineup to one-tenth of the original price. But I do not think pricing is the most interesting part anymore, because there are already many pieces making that point. I think the more interesting point is what V4 says about China’s open-source AI system and how it is slowly justifying the question that is asked the most: why?
Open source is not just an ideology here. It is an R&D efficiency mechanism. Because Chinese labs are compute-constrained, capital-constrained, and increasingly trying to build around domestic hardware, they cannot afford to waste as many resources duplicating the same low-level/infra-level research. If one lab solves an architecture problem, an inference optimization problem, or a Huawei adaptation problem, the rest of the ecosystem can absorb it. Even if they do not necessarily get the full training recipe, data mix, serving kernels, or CANN-specific optimization code, they still get a valuable reference point.
In that sense, DeepSeek is not just competing with the other Chinese labs. It is also providing shared infrastructure for them, whether out of the kindness of its heart or from some pressure from up top.
Jensen Huang has been talking about AI as a five-layer cake: energy, chips, infrastructure, models, and applications, and reiterating the importance of being plugged into China’s ecosystem. Nvidia’s own framing is that every successful application pulls on every layer beneath it, from models all the way down to energy. And with the recent push of its open-source models, it is again signaling to the industry that open-source models will benefit the ecosystem. This seems to keep getting lost in translation. The usual interpretation is about how Nvidia touches multiple layers of that stack and thus can sell more of its compute, but I think it’s beyond that.
What V4 shows is that in the China AI sector, even inside just one layer — the model layer — the labs are already starting to behave like a more coordinated system. The model layer itself is becoming a kind of shared R&D layer that helps propel the rest of the stack forward.
That is the frame I want to use for V4. Not “DeepSeek is cheap, therefore everyone else gets hurt,” although that is partly true. The more interesting story here is that DeepSeek is pushing base-layer innovation into the open, and the rest of the Chinese AI ecosystem is reorganizing around that.
TLDR, my key arguments:
1/ Chinese AI labs are starting to look like specialized units inside one compute-constrained mega-lab. DeepSeek does the base infrastructure work; others specialize in coding, agentic workflows, multimodality, distribution, and product integration.
2/ Open source can still make money. Users are not just paying for model weights. They are paying for managed inference, reliability, latency, maintenance, routing, memory, tools, and the whole agent harness. Just like how open-source software before made money.
3/ The bigger role of open-weighted releases is shared R&D. If one lab solves an architecture, inference, or hardware-adaptation problem, the rest of the ecosystem can study the released model, benchmark it, distill from it, and optimize around it instead of wasting scarce compute rediscovering every piece from scratch.
4/ The Huawei adaptation should be viewed through this lens. Making V4 run well on Huawei’s stack took real R&D effort, even if it’s only just inference first. The read I have heard from Chinese AI circles is that DeepSeek delayed V4 by almost three months to make this work properly. If true, DeepSeek basically took one for the team.
5/ The pricing squeeze on Z.ai, MiniMax, and Kimi is real, but secondary. It may also be temporary, because these labs can incorporate DeepSeek’s innovations and lower their own cost structure over time.
Open source can still make money
There is a common question around China’s open-source / weight AI ecosystem: how do you make money if you open-source the model?
The answer to this a year ago was that there is no pressure. Then the pressure came, and the answer was through other functions through superapps. Then the thinking was to sell cloud and then APIs. But now it finally hit me, I think the answer is pretty simple. Open-weight does not mean free inference. It means zero model-provider take rate if the user self-hosts, but how many users are realllllyyy self-hosting?
Self-hosting a frontier-scale model is not easy. Users still need hardware, power, memory, depreciation, utilization, maintenance, upgrades, latency optimization, reliability, and the whole inference stack. And as AI becomes more agentic, the harness around the model becomes even more important: routing, tools, memory, retrieval, workflow orchestration, monitoring, security, all of it.
If you’re like me, you probably asked, ‘If it’s all open, then why are there quoted variations of API prices?’ BECAUSE most users are not really paying for access to model weights. They are paying for managed inference.
This is basically the open-source software model. The code can be open, but the managed service is where the money is. That is why open-source model companies can still monetize. They can still charge for API usage. They can still sell enterprise deployment. They can still sell private cloud, reliability, support, workflow integration, and model customization.
But open source changes the ceiling on monetization. If users have a strong open model available, and if sophisticated buyers can self-host or compare against a self-hosting cost floor, then the model provider cannot charge unlimited premium pricing for generic inference. So open source does not destroy monetization. It disciplines it.
That is the right way to think about V4 and what it does to the likes of z.ai, Minimax, and Moonshot. Reuters reported that DeepSeek V4’s models are available as open-source releases under a permissive MIT license, allowing companies to freely use, modify, and commercialize them, thus making the engineering behind V4 more valuable to everyone else.
Open source as shared R&D
But the more important point is not monetization. It is R&D efficiency for a whole ecosystem.
If every Chinese lab had to independently rediscover the same architecture tricks, inference optimizations, memory improvements, and hardware adaptations, that would be a huge waste of scarce capital, talent, and compute. The more efficient equilibrium is that one lab makes a breakthrough, releases it, and the rest of the system absorbs it.
This is what China’s open-source ecosystem is starting to look like. Even Open-weight models can still function as shared R&D, but the transfer is not total. Other labs get weights, architecture clues, benchmark behavior, technical reports, deployment experience, and a reference model to study.
When DeepSeek releases V4, other labs do not just get a competitor. They also get a base to study. They can look at the architecture, copy what works, improve their own inference efficiency, and spend less time repeating the same base-layer experiments. That lets them spend more time on higher-level product behavior: coding reliability, agent orchestration, workflow integration, multimodal usefulness, enterprise deployment, and post-training.
This is why I think DeepSeek is becoming the shared foundation-layer R&D engine for Chinese AI. Just as it was implied here before.
This is something I have heard directly from people working at Chinese labs. They describe it as shifting more R&D resources toward “higher-level” innovations — coding behavior, reinforcement learning, agent orchestration, workflow reliability, product-specific usefulness — while waiting for DeepSeek to open-source more of the “lower-level” infra-layer innovations.
As an analogy, imagine OpenAI, Anthropic, and Google DeepMind each focusing on one part of the research stack and then sharing their outputs with one another. That is not how the US system works. But it is, to some degree, how the Chinese system is evolving.
And we go back to “necessity is the mother of innovation.” China has less access to frontier Nvidia chips. Compute is scarce. Capital is not infinite. Talent is limited (labs are 1/10 of the headcount compared to leading labs in the US). So the ecosystem has less (no) room to waste. Open source becomes a coordination mechanism. And DeepSeek, whether purposely or not, is betting on the fact that the rest of the model companies will incorporate its engineering innovation in their next iterations.
China’s labs are specialized units
If you take a step back and look at each and every one of the labs, they look like they’re each pursuing a different strategy these days. This could be due to skewed talent, founder taste, or pressure to show differentiation as they’ve gone public. But put together, they’re like a mega lab, almost like each a business unit.
Again, these companies are not literally one company. They are still competitors. But under the open-source layer, they increasingly look like specialized units within a single large, compute-constrained R&D system.
DeepSeek is doing foundation-layer architecture, efficiency, and now Huawei-stack adaptation. Z.ai is leaning into coding and enterprise reliability. Coding is one of the few AI workloads where the process is structured enough for agents to matter, and it is one of the clearest enterprise wedges. Kimi is leaning into long context and agentic orchestration. The Kimi product experience has always been more about getting the model to handle long, messy, multi-step tasks in a useful way. MiniMax is leaning into multimodality, voice, video, and low-cost inference. It is not trying to win only by having the best text model. Alibaba and Tencent are leaning into distribution, cloud, workflow, and product integration. ByteDance is leaning into consumer distribution, Douyin, Feishu, CapCut, and short-video workflows.
So the Chinese ecosystem is not just “a bunch of model labs.” It is a set of labs, each trying to own a different part of the stack. And because much of the base model layer is open or semi-open, each lab can absorb what the others discover and then specialize further.
This feels very different from the US system. In the US, OpenAI, Anthropic, and Google each try to own as much of the full stack as possible. They each run their own experiments, hit their own dead ends, optimize their own hardware relationships, build their own inference stack, and keep most of the output private. That structure has advantages. It allows fast private compounding. It protects the research frontier. It lets the winning lab capture more value. But it also means a lot of repeated work.
China does not have the same luxury, if you must, so the system is becoming more modular. DeepSeek pushes the base infrastructure forward. Other labs build on top.
I think that is the bigger picture that is being overlooked. When I started this piece, my first thought was that DeepSeek hurts its peers commercially by squeezing their Anthropic-like tiered performance/ usage pricing models. But as I processed the information, I realized that’s only the story in the short term; rather, in the longer term, DeepSeek will once again propel its technical capabilities, and the pricing model will readjust.
Huawei adaptation is also shared R&D
Now, if you view the Huawei part of this story through the same lens. We can see that making a near frontier model run well on Huawei’s stack is not just a political move or a ragey response to export controls, nor is it a clickbait press release. It takes real engineering work. It means dealing with CANN, kernels, memory layout, numerical formats, inference bottlenecks, latency, throughput, compiler issues, and all the tedious software-hardware co-optimization work that developers normally prefer to avoid if Nvidia is available. This is what Jensen was saying: that most developers want to use CUDA and that it took them 10+ years to build that kind of stickiness out of habit, due to quality of service, and so on. It’s pretty safe to say that most developers think running on Nvidia is just easier.
So, the read I have heard from people in Chinese AI circles is that DeepSeek delayed V4 by almost three months to ensure the model would run properly on Huawei’s stack before launch. If true, that is a very important signal.
Most other Chinese labs do not have the luxury to do this. They need to launch, show progress, monetize, raise money, keep customers engaged, and prove that their model is still competitive. If they have a strong model ready, the rational commercial move is to ship it. Spending three to six extra months trying or even attempting to switch to Huawei adaptation is too costly.
In this sense, DeepSeek kind of took one for the team. It absorbed the painful adaptation work, then released a model that gives the rest of the Chinese ecosystem a reference workload for Huawei. That is not just good for DeepSeek. It is good for Huawei, for CANN, for Chinese cloud providers, and eventually for every Chinese lab that wants to reduce Nvidia dependency.
As we know, a key change from earlier DeepSeek releases is that V4 inference was adapted for Huawei’s most advanced Ascend AI chips. It is also reported that Huawei said V4 is fully supported on its Ascend 950-based supernode clusters, and that the entire Ascend supernode product line now supports the DeepSeek-V4 series models.
So this is where it goes beyond the model layer and crosses to the next layer in the cake, the hardware layer. In Jensen’s cake metaphor, the model layer in China is now actively pulling the chip and infrastructure layers forward. DeepSeek is not just releasing weights. It is helping create the workload around which the domestic hardware stack can improve. This again explains why open-source? Now they can all try it out.
The pricing squeeze is real, but not the main story
As I said, this piece was initially 80% about the pricing squeeze, but then I realized that it is only temporary. But this does not mean the pricing angle is wrong.
V4 still pressures Z.ai, MiniMax, and Kimi, especially in text, coding, and agentic workloads. Any lab trying to charge a premium for generic reasoning or coding now has to explain why customers should not route more volume to DeepSeek, or at least use DeepSeek as the reference price. Before V4, the Chinese frontier labs were trying to move pricing higher - quite literally, all of them did it. Compute was tight, demand was growing, and investors wanted to see monetization. Some raised prices to make more, some used price as a way to filter out users because they could not serve as many as they’d like to. Not to be too repetitive, but because V4 is open, other labs can study its architecture, adopt its inference optimizations, and reduce their own serving costs over the next few months. So DeepSeek is compressing peers’ pricing today but also giving them the technical path to lower their own costs tomorrow, and raising the ecosystem’s technical floor.
Furthermore, the question is not only whether V4 beats Kimi, GLM, or Qwen on a specific benchmark. If V4 has better inference efficiency, the rest of the ecosystem can learn from it. If V4 has better agentic behavior, the rest of the ecosystem can post-train around it. If V4 is optimized for Huawei, the rest of the ecosystem now has a stronger reference point for domestic deployment.
The US contrast
This kind of partial switch from CUDA to CANN takes serious effort. For example, Anthropic has spent serious effort making Claude work across multiple compute platforms. It announced an expanded Google/Broadcom partnership for next-generation TPU capacity, and it has separately described deep technical collaboration with AWS on Trainium, including writing low-level kernels and contributing to the AWS Neuron software stack. Amazon also said Anthropic will secure up to 5GW of current and future Trainium capacity, and that Amazon and Anthropic engineers communicate on everything from low-level optimization work to high-level architectural decisions for next-generation chips.
That is strategically rational for Anthropic. The difference here is that no one would expect Anthropic’s Trainium / TPU optimization work to become shared infrastructure for OpenAI.
In the US right now, hardware optimization work becomes part of each lab’s private moat. In China, DeepSeek is doing the Huawei adaptation to not just benefit DeepSeek. It potentially becomes a reference point or even a standard for the rest of the ecosystem.
To be clear, I am not saying China’s structure is automatically better. The US system still has deeper compute, better chips, more capital, more mature cloud infrastructure, and the leading closed frontier labs. But China’s system may be more R&D-efficient per unit of scarce compute. And it is working more closely like a team, and again, under constraint, that can make a huge difference.
Why DeepSeek can behave this way
Most independent labs cannot behave this way, obviously, nor are they that just ‘kind-hearted.’ They need to launch models, acquire users, raise capital, show revenue, and monetize API demand. If they have a working model, the rational move is to ship it.
DeepSeek seems to have more room to optimize for ecosystem strategy rather than immediate monetization. That is why it can delay a release for Huawei adaptation, publish foundation-layer improvements, keep API prices low, and still operate as if the broader ecosystem benefit matters.
We’ve written about how the founder is philosophically quite committed to open-sourcing frontier technology, but in many ways, we can interpret DeepSeek as becoming somewhat of a national strategic asset in function, even if not formally nationalized in ownership. Since it’s backed by High Flyer, the quant hedge fund that is no short of dough, and it is said to continue to serve large enterprise clients. It is under no pressure to make a quick buck now, with its recent round of valuation at ~20billion USD, it is absorbing some of the infrastructure burden for the whole system.
Big tech benefits from this
Now, how do the big tech players fit into this? Tbh they’re kinda in the background and not that important for this strategic piece, except for the fact that they’re all financially linked to the labs (rumors are Baba and Tencent will invest in DeepSeek’s next round). What they are good at is that they are structurally better positioned to diffuse AI than standalone labs.
Alibaba, Tencent, and ByteDance do not need to monetize only through tokens. They can use cheaper models to improve cloud, commerce, ads, payments, enterprise software, productivity tools, games, short video, search, browsers, mini programs, and developer tools. If DeepSeek makes the model layer cheaper, that is not necessarily bad for them. It can actually be useful.
They can integrate whichever model works best at the time. Tencent did this with DeepSeek before, and it can do it again. One caveat is that Tencent’s latest HY3 preview was released in the same week as DeepSeek, and the limelight seems to be completely stolen.
Independent labs are in a harder position. Z.ai, Moonshot, and MiniMax have to compete more directly on model performance, coding, agents, hallucination rates, private deployment, or specific consumer usage. They do not have the same built-in cloud, payments, commerce, social graph, or short-video distribution. So that makes them more exposed to model deflation.
The duality of DeepSeek
So the real lesson from V4 isn't just that DeepSeek is cheap. It hurts other Chinese labs as competitors, but helps them with infrastructure. I think the bigger story here is that China’s open-source AI ecosystem is becoming increasingly coordinated, and each has a role to play in propelling the whole sector forward, given compute constraints, in efforts to optimize R&D.
DeepSeek pushes the base infrastructure forward. Other labs absorb the work and specialize further up the stack. Huawei gets a serious reference workload. Standalone labs face pricing pressure today, but also have a path to lower their own costs tomorrow.
Anyway, Part 2 is about V4’s implications on Huawei and Nvidia and what this symbolizes across the AI stack (the five-layered cake).
Some relevant reads that helped me in understanding the pricing economics better and DS analysis:
DeepSeek V4 preview release. DeepSeek’s official release page on V4-Pro and V4-Flash, open weights, parameter counts, and 1M context.
South China Morning Post. Huawei and DeepSeek collaboration, day-zero Ascend adaptation, and DeepSeek’s throughput caveat until Ascend 950PR ships at scale.
Reuters. DeepSeek V4 adapted for Huawei chips, V4-Pro and V4-Flash positioning, Huawei support on Ascend 950 clusters, and Huawei’s claim that its chips were used for part of V4-Flash training.
Bernstein, “China Internet: A primer on AI token economics and inference margins,” Robin Zhu et al., 21 April 2026. Q1 2026 compute price hikes, GLM-5 vs M2.5 cache pricing and hit-rate data, Qwen 3.5 cost-vs-intelligence trajectory, Alibaba Token Hub, “low-cost agentic back-end” commoditisation framing.
Bernstein, “China Internet: AI inference margins addendum; DeepSeek-V4, HY3-preview thoughts,” 27 April 2026. V4-Pro and V4-Flash blended pricing, 75 percent discount within 24 hours of launch, MiniMax throughput reconciliation, Tencent Muse Spark framing, DeepSeek fundraising base case.
Jukan / Citrini, X thread. The argument that V4’s architectural pattern aligns with Nvidia’s Rubin and G3.5 roadmap rather than simply eroding it.
Jensen Huang on the Dwarkesh Patel podcast. The argument that US export controls accelerate the development of a parallel Chinese chip ecosystem.
Tencent HY3 preview official release. HY3 architecture, open-source release, Tencent ecosystem integrations, CodeBuddy / WorkBuddy improvements, TokenHub pricing, and inference efficiency.




Grace is the GOAT 🐐
Feels like an appropriate place to use 工合, or "gung ho" in the original sense of the word.. maybe even 人工合作 if some native Chinese speaker hasn't made a snappier term.