DeepSeek open-sources a full software stack for Huawei's Ascend chips, taking aim at Nvidia's CUDA

DeepSeek open-sourced a full suite of infrastructure software for Huawei's Ascend AI chips on September 30, 2026, releasing six modules that together cover programming, compute optimization, and distributed communication — the same layers of tooling that have made Nvidia's CUDA ecosystem so difficult for competitors to dislodge.
At the center of the release is an Ascend-specific version of TileLang, a programming language originally developed by researchers at Peking University that DeepSeek has used internally for roughly a year. TileLang is designed to let developers write AI kernels at a higher level of abstraction while still hitting hardware-level performance, abstracting away much of the low-level code that chip programming typically requires. DeepSeek has explicitly positioned TileLang as offering “a simpler programming model” than CUDA — a direct challenge to the interface that has anchored Nvidia's software moat for over a decade.
What's actually in the release
Alongside TileLang, DeepSeek published DeepGEMM and DeepEP for optimized matrix computation, TileKernels and FlashMLA for compute workloads, and DeepSelect for workload routing — libraries that handle both the heavy-lifting math operations AI models require and the communication overhead of running them across many chips at once. The two companies also detailed a joint “supernode” reference design built on 128 Ascend 950 chips, intended to demonstrate that Huawei's hardware can be optimized for both computation and inter-chip communication at data-center scale.
Why this matters beyond one company's tooling choices
Nvidia's dominance in AI compute rests as much on software lock-in as on raw chip performance — CUDA's maturity, library ecosystem, and years of developer familiarity make switching hardware platforms costly even when competing silicon is technically capable. By open-sourcing a comparably deep software stack for Ascend, DeepSeek is attacking that lock-in directly rather than just competing on hardware specifications.
The move also reflects a broader pattern of Chinese AI companies consolidating around domestic hardware as US export controls continue to restrict access to Nvidia's most advanced chips. An open-source, Chinese-developed software layer removes one of the biggest practical barriers — tooling maturity — that has kept Ascend chips a harder sell than Nvidia GPUs even among Chinese developers who would otherwise prefer to avoid export-control exposure.
Whether TileLang and its accompanying libraries actually narrow the performance gap with CUDA in practice remains to be tested at scale by developers outside DeepSeek and Huawei. But the release itself signals that China's AI ecosystem is treating software infrastructure, not just chip fabrication, as a strategic priority worth open-sourcing rather than keeping proprietary.
Originally reported by The Decoder. Read the original article for additional details.
View original source