i
News
News · 2026-09-30

DeepSeek and Huawei target Nvidia’s software advantage with Ascend tools

@neuronium_ai @neuronium_ai

DeepSeek and Huawei have released open-source software for Huawei’s Ascend AI chips, including libraries for computation and data transfer between chips. The companies also optimized a supernode built from 128 Ascend 950 chips. The release targets a less visible constraint on China’s chip ambitions: hardware matters little if developers cannot make it run efficiently. At its center is TileLang, a programming language DeepSeek says is easier to use than CUDA.

Cover: DeepSeek and Huawei target Nvidia’s software advantage with Ascend tools

The software gap

DeepSeek says Huawei fully supported the work. TileLang was created by researchers at Peking University, and DeepSeek has used it for about a year. The company argues that an independent software ecosystem for AI chips needs a general-purpose language that is both easy to use and capable of extracting the hardware’s full performance.

DeepSeek first tested TileLang on older Nvidia chips. It is now the company’s main tool in its work on artificial general intelligence, The New York Times reported.

Nvidia’s position rests on more than chip design. About four million developers worldwide work with CUDA, according to an estimate cited in the Russian report. That installed base has been a barrier for competitors such as AMD, even when their hardware looked comparable on paper.

Huawei announced new AI processors and supernode systems two weeks before DeepSeek’s announcement, saying they would be widely used for model training next year.
Huawei says it cannot meet domestic demand and plans to reduce chip sales abroad.
Huawei’s rotating chairman, Eric Xu, said the company could not accept a future in which China’s access to chips depends on other countries being willing to sell them.
1New Ascend processors
2TileLang tools
3128-chip supernode

What CUDA’s lead does — and doesn’t — prove

SemiAnalysis recently tested Jalapeño, OpenAI’s inference chip, and called CUDA’s advantage “possibly already gone”: OpenAI can bring new models up quickly on its own hardware. In most tested scenarios, Jalapeño delivered better performance per watt than Nvidia’s Blackwell. SemiAnalysis said OpenAI’s models helped design the chip, even though those models themselves run on Nvidia GPUs.

That result has limits. The analysts tested only workloads they considered relatively easy to optimize: about 8,000 input tokens and 1,000 output tokens. They had not yet run AgentX, a benchmark for how AI agents handle multistep tasks.

In August, SemiAnalysis found Nvidia had a substantial advantage on those tasks. With AMD’s current software stack, Nvidia would still be cheaper per token even if AMD gave its hardware away for free. The analysts attributed Nvidia’s long-term edge not to the chips themselves, but to the software that connects many chips into one system.

Huawei’s chips were not tested on AgentX. In an earlier analysis of DeepSeek V4, however, SemiAnalysis said Huawei’s CANN software stack was the only one besides CUDA to support the model on its first day.

I think the key test for DeepSeek and Huawei is not whether TileLang makes an individual chip easier to program. It is whether the software can keep pace as workloads grow more complex and require many chips to work together. The announcement gives a concrete start — libraries, a language and a 128-chip system — but says little about performance on the demanding agent workloads where Nvidia’s software advantage has mattered most.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X