- China’s leading AI model developers are eyeing custom inference chips for one reason: economics, not prestige
- The real threat to Nvidia isn’t another GPU. It’s the possibility that customers finally have an alternative
In July, reports that DeepSeek was developing its own AI inference chip gathered momentum.
Reuters, citing people familiar with the matter, reported on July 7 that the project began about a year ago.
DeepSeek has reportedly approached chip design firms, foundries and memory suppliers, while quietly recruiting chip engineers through private channels.
The market reaction was immediate: Nvidia shares fell about 1.6% in pre-market trading.
Around the same time, Chinese AI startup Zhipu was also reported to be evaluating custom AI chips.
The timing was unlikely to be coincidental.
China’s two leading foundation model developers are both looking beyond algorithms and toward silicon. At first glance, the move looks like a declaration of technological ambition. In reality, it is more likely a carefully calculated business decision.

The biggest significance of DeepSeek’s chip effort may not be replacing Nvidia. It may be weakening the very logic that has allowed Nvidia to maintain extraordinary margins for years: customers have no alternative.
Why inference chips?
Designing chips for AI training is one of the most difficult engineering challenges in computing.
Training clusters containing tens of thousands of GPUs must operate continuously for months, where a single failure can wipe out millions of dollars’ worth of computing time.
More importantly, Nvidia has spent two decades building CUDA into the industry’s default software ecosystem. Millions of developers, frameworks and AI tools are deeply tied to it. That ecosystem—not just the hardware—is Nvidia’s true moat.
Inference presents a different problem.
Unlike training, inference simply runs an already-trained model. The computation graph is relatively fixed, making it easier to optimize hardware for a specific workload instead of supporting every possible AI application.
Google demonstrated this years ago with the first-generation TPU. Built on a relatively modest 28-nanometer process, TPU v1 delivered roughly 15 to 30 times better inference performance than contemporary CPUs and GPUs for certain workloads.
That is why inference chips represent a much more realistic target than challenging Nvidia head-on in training GPUs.
Every token has a cost
Training is a one-off capital investment. Inference is a utility bill that arrives every month.
Once a model serves trillions of tokens each week, shaving even a tiny fraction of a cent from the cost of each token can translate into hundreds of millions of yuan in annual savings.
Purpose-built inference chips eliminate much of the unnecessary flexibility required by general-purpose GPUs, allowing higher energy efficiency and lower operating costs.
TrendForce expects the AI ASIC market to grow 44.6% in 2026, far outpacing the projected 16.1% growth for GPUs. ASIC stands for application-specific integrated circuit.
For companies running massive AI services, reducing inference costs is no longer merely an engineering optimization—it is becoming a strategic imperative.
‘Changing engines mid-flight’
DeepSeek’s journey toward greater hardware independence has hardly been smooth.
Industry sources say that during the middle of 2025, DeepSeek encountered repeated stability problems while training its V4 model on Huawei’s Ascend chips.

Training runs reportedly crashed midway, inter-chip communication fell short of expectations, and overall system stability proved insufficient.
The problems delayed V4’s launch from its original target around the 2026 Lunar New Year until late April.
The delay reflected what engineers often describe as changing an aircraft’s engines while it is still in flight.
According to Chinese media reports, migrating from Nvidia’s CUDA ecosystem to Huawei’s CANN framework required rewriting more than 200 core operators, representing roughly 30 person-years of low-level engineering work.
Early versions achieved only about one thirty-fifth of the inference performance eventually reached after optimization.
After months of engineering effort, inference speed on Huawei’s Ascend 950PR reportedly improved by a factor of 35.
The experience also reinforced a broader lesson. When general-purpose GPUs cannot perfectly match a model’s workload, the logical next step is to design hardware around the model itself.
Beyond the chip
Nvidia has maintained gross margins above 75% largely because buyers have had nowhere else to go for cutting-edge AI accelerators.
Having no alternative is itself a form of pricing power.
DeepSeek’s custom chip initiative is unlikely to threaten Nvidia’s technological leadership anytime soon. But it may reshape negotiations.
The moment a procurement manager can tell Nvidia that its own inference chip is expected to tape out next year, the conversation changes. Even if that chip is not yet commercially ready, the mere existence of a credible alternative weakens Nvidia’s leverage.
The significance of DeepSeek’s effort lies less in the chip itself than in what it represents.
As model architectures gradually converge and algorithmic improvements become harder to achieve, hardware is emerging as the next battleground for differentiation.
Leading AI companies increasingly recognize that they cannot leave control of computing costs—and pricing power—entirely in someone else’s hands.



