DeepSeek has open-sourced a full programming toolchain for Huawei's Ascend AI accelerators, built around TileLang, a tile-based kernel language it pitches as a simpler alternative to Nvidia's CUDA. The announcement matters less as a chip story than as a software one: CUDA's lock-in has always lived in its libraries and developer habits, and DeepSeek just published Ascend versions of the libraries its own models depend on.
- On Sept 30, 2026, DeepSeek announced via WeChat a free suite of tools for Ascend, developed with Huawei's full support.
- The core is a version of TileLang, which DeepSeek now calls its "core tool," plus six compute and communication modules that mirror its earlier Nvidia releases.
- It targets Huawei's Ascend 950 chips, including a "supernode" of 128 accelerators, and DeepSeek plans at least 160,000 Huawei accelerators at an Inner Mongolia data center.
- The gaps are real: tooling maturity, profilers and PyTorch operator coverage on Ascend still trail the CUDA world.
What did DeepSeek actually release?
Three layers. At the bottom is the Ascend adapter for TileLang, published on GitHub as tile-ai/tilelang-ascend. TileLang is a kernel language where you reason about tiles of data instead of individual threads, and DeepSeek says it was first tested on older Nvidia chips before the Ascend port. Above it sit six software modules covering compute and communication, each mirroring something DeepSeek previously shipped for Nvidia GPUs. Coverage names DeepGEMM for matrix multiplication, FlashMLA for multi-head latent attention kernels, TileKernel, DeepSelect and DeepEP for expert-parallel communication.
RelatedPython 3.14 makes free-threading real: the GIL era is ending
Two of those are easy to verify. The deepseek-ai/DeepGEMM-Ascend repository describes itself as a clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs and was created on Sept 29. deepseek-ai/DeepEP-Ascend, a high-performance communication library for training and inference on Ascend NPUs, followed on Sept 30. There is also an older clangd-ascend repo from May 2026 for code lint and completion, a small hint that this was a months-long effort rather than a weekend port.
The framing is deliberate. In DeepSeek's words: "To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential."
Why is the software the real story?
Ask anyone who has tried to move a serious training stack off Nvidia and they will not start with the chips. They will talk about the hand-tuned kernels, the communication primitives, the profilers, and the muscle memory of roughly 4 million CUDA developers, a figure cited in coverage of the release. The silicon can be matched on paper. The accumulated software cannot be copied in a quarter.
A tile-level language attacks that cost directly. When kernels are written against tiles and the compiler handles the mapping to a specific backend, adding a new target means writing a backend, not rewriting a library. That is the bet behind tilelang-ascend. And because DeepSeek published Ascend versions of the same named libraries, models built on its stack can move across with far less friction than a typical port.
| Layer | DeepSeek Ascend stack | CUDA stack |
|---|---|---|
| Kernel language | TileLang (tilelang-ascend) | CUDA C++ / PTX |
| Matrix multiplication | DeepGEMM-Ascend | DeepGEMM (Nvidia original) |
| Attention kernels | FlashMLA-style MLA kernels | FlashMLA (Nvidia original) |
| Communication | DeepEP-Ascend | DeepEP (Nvidia original) |
| Developer base | Brand new, open source | About 4 million developers |
| Hardware scale | 128-chip Ascend 950 supernode; 160,000 accelerators planned | Nvidia GPUs, years of production deployment |
How far does the Nvidia originals' popularity carry over?
The Nvidia versions of these libraries are not obscure. FlashMLA has about 13,000 GitHub stars, DeepEP about 10,200 and DeepGEMM about 7,900. Those are the developers who already read, fork and benchmark DeepSeek's kernels, and the Ascend ports give them a familiar API surface to point at different hardware. Stars are a crude proxy, but they describe the audience the ports inherit on day one.
What are the limits?
Publishing a library is not the same as matching its performance, and nothing in the announcement includes head-to-head benchmarks against the CUDA versions. An arXiv field study (2607.08215) on the limitations of non-GPU accelerators for mixture-of-experts and multimodal serving on Ascend is a useful reality check: maturity, profiling tools and breadth of PyTorch operator coverage are still where the friction lives. CANN, Huawei's own stack, remains underneath all of this, so DeepSeek is building on a base it does not control.
RelatedNixpkgs core team disbands, citing steering committee friction
Hardware economics add a second constraint. Our earlier reporting found Huawei raising Ascend 950DT prices by 20 to 50 percent in two months on grey-market HBM costs, and Huawei plans to restrict chip sales outside China. A better software stack does not fix a scarce memory supply.
What it means for the market
The signal for investors is slow rather than sudden. Nvidia (NVDA) China revenue is already shaped by export controls, and a credible domestic software path erodes the argument that Chinese labs have no alternative once they get the chips. That is a story measured in years, not quarters. For Cambricon (688256.SS) and SMIC, the read-through is demand: the more a flagship lab standardizes on domestic accelerators, the more the supply chain behind them matters. Huawei is private, so the exposure runs through its suppliers and rivals. DeepSeek itself reached a $1 billion annualized revenue run rate in late September, with a valuation near 500 billion yuan (about $74 billion), so it has the weight to make an ecosystem choice stick.
- Sep 11Ascend 950DT price rises. Up 20 to 50 percent in two months on HBM costs.
- Sep 17960DT pulled forward. Huawei moves it to Q1 2027 on a yearly cadence.
- Sep 30DeepSeek opens its Ascend toolchain. TileLang adapter plus six libraries.
- Q1 2027Ascend 960DT arrives.
- Benchmarks. Ascend kernels against their CUDA counterparts on the same models.
- Framework adoption. Whether PyTorch and vLLM users start targeting tilelang-ascend.
- The 960DT. Huawei's Q1 2027 part is the next test of the hardware cadence.
- DeepSeek's next model. One trained and served on Ascend would turn this from tooling into proof.
Our take
This is the most credible attack on CUDA we have seen precisely because it is boring. No new chip, no benchmark victory lap, just the unglamorous libraries a lab needs to run its own models. Nvidia's lead was never only about transistors, and DeepSeek has picked the layer where a determined, well-funded team can actually make progress. It will not unseat CUDA, whose developer base is orders of magnitude larger. What it can do is make Ascend a workable second target for the one customer base Nvidia can least afford to lose. Whether it works will show up in benchmarks and in who ports next, not in the announcement.
- GitHubtile-ai/tilelang-ascend the Ascend TileLang adapter
- GitHubdeepseek-ai/DeepGEMM-Ascend matrix multiplication kernels for Ascend NPUs
- GitHubdeepseek-ai/DeepEP-Ascend communication library for Ascend NPUs
- NewsThe Next Web coverage of the release
- PaperarXiv 2607.08215 limits of non-GPU accelerators for MoE serving
Original analysis by GenZTech. Reporting via The Next Web.
