The software gap
DeepSeek says Huawei fully supported the work. TileLang was created by researchers at Peking University, and DeepSeek has used it for about a year. The company argues that an independent software ecosystem for AI chips needs a general-purpose language that is both easy to use and capable of extracting the hardware’s full performance.
DeepSeek first tested TileLang on older Nvidia chips. It is now the company’s main tool in its work on artificial general intelligence, The New York Times reported.
Nvidia’s position rests on more than chip design. About four million developers worldwide work with CUDA, according to an estimate cited in the Russian report. That installed base has been a barrier for competitors such as AMD, even when their hardware looked comparable on paper.
What CUDA’s lead does — and doesn’t — prove
SemiAnalysis recently tested Jalapeño, OpenAI’s inference chip, and called CUDA’s advantage “possibly already gone”: OpenAI can bring new models up quickly on its own hardware. In most tested scenarios, Jalapeño delivered better performance per watt than Nvidia’s Blackwell. SemiAnalysis said OpenAI’s models helped design the chip, even though those models themselves run on Nvidia GPUs.
That result has limits. The analysts tested only workloads they considered relatively easy to optimize: about 8,000 input tokens and 1,000 output tokens. They had not yet run AgentX, a benchmark for how AI agents handle multistep tasks.
In August, SemiAnalysis found Nvidia had a substantial advantage on those tasks. With AMD’s current software stack, Nvidia would still be cheaper per token even if AMD gave its hardware away for free. The analysts attributed Nvidia’s long-term edge not to the chips themselves, but to the software that connects many chips into one system.
Huawei’s chips were not tested on AgentX. In an earlier analysis of DeepSeek V4, however, SemiAnalysis said Huawei’s CANN software stack was the only one besides CUDA to support the model on its first day.
I think the key test for DeepSeek and Huawei is not whether TileLang makes an individual chip easier to program. It is whether the software can keep pace as workloads grow more complex and require many chips to work together. The announcement gives a concrete start — libraries, a language and a 128-chip system — but says little about performance on the demanding agent workloads where Nvidia’s software advantage has mattered most.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X