Deepseek is releasing open-source programming tools for Huawei AI chips. The centerpiece is TileLang, a language meant to be easier to work with than Nvidia's CUDA software platform.
Chinese AI developer Deepseek has teamed up with Huawei to build programming tools for Huawei's Ascend chips, according to a post on Deepseek's official WeChat channel. The software includes libraries for computation and for moving data between chips, and Deepseek is making all of it open source, Reuters reports. According to Deepseek, Huawei fully supported the work. The two companies also optimized a so-called supernode, a cluster of 128 Ascend 950 chips.
TileLang is at the core of the release. The open-source programming language for AI chips was originally developed by researchers at Peking University, and Deepseek has been using it for about a year. Deepseek argues that anyone trying to build an independent software ecosystem for AI chips first needs a universal language that's easy to program but still gets full performance out of the hardware. In the company's view, TileLang offers a simpler programming model than CUDA. Deepseek first tested the language on older Nvidia chips. TileLang is now the company's main tool for its work on artificial general intelligence, The New York Times reports.
Read nextOpenAI Releases GPT-6.1 Sol as Flagship Astra Stalls on SafetyChina's AI industry faces a software gap
The partnership goes after one of the biggest problems facing China's AI industry. Domestic chips need software that can get the most out of them. Nvidia's dominance doesn't come from chip design alone. It also rests on an estimated four million developers worldwide who build with CUDA. That ecosystem is the moat rivals like AMD haven't been able to cross, even when their hardware looked just as strong on paper. So far, Chinese model makers like Z.ai and Moonshot AI have moved faster than the country's chipmakers, according to the NYT.
Huawei wants to close that gap. Two weeks before Deepseek's announcement, the company unveiled new AI processors and supernode systems and said they would be widely used for model training next year. Huawei also admits it can't keep up with demand at home, so it plans to sell fewer chips abroad. Referring to US export controls, Huawei's current rotating chairman Eric Xu said the company can't accept a future that hinges on whether others are willing to sell chips to China.

Analysts say the CUDA moat is potentially dead
Research firm SemiAnalysis has been looking at how much of that moat is left. After testing Jalapeño, OpenAI's inference chip, the analysts called the CUDA moat potentially dead because OpenAI gets new models running on its own hardware so quickly. Jalapeño beat Nvidia's Blackwell on performance per watt in most of the scenarios tested. According to SemiAnalysis, OpenAI models also helped design the chip, and those models run on Nvidia GPUs.
The analysts added their own caveats. They only tested scenarios that are relatively easy to optimize, with about 8,000 input tokens and 1,000 output tokens. They haven't yet run AgentX, a benchmark that measures how AI agents handle multistep tasks. That's exactly where SemiAnalysis found Nvidia well ahead back in August. With AMD's current software stack, Nvidia would still come out cheaper per token even if AMD gave its hardware away. The authors don't see Nvidia's lasting advantage in the silicon itself. They see it in the software that links many chips into one system.
Huawei's chips weren't part of the AgentX comparison. In an earlier analysis of DeepSeek V4, however, SemiAnalysis pointed out that Huawei's CANN software stack was the only one besides CUDA to support the model on day one.



