The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Compute & Chips · DeepSeek · Huawei · Ascend · TileLang · DeepGEMM · FlashMLA

DeepSeek open‑sources its kernel libraries for Huawei's Ascend chips

DeepSeek says every TileLang operator used in training its V4 models now has an Ascend version. FlashMLA's README claims 95% of the Ascend 950's peak on prefill.

DeepSeek has open-sourced infrastructure components for Huawei's Ascend chips, according to the account MaxForAI, which listed TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect. The researcher eliebakouch said DeepSeek updated most of its open-source libraries with Ascend support. Both posts went up around 02:45 UTC on 30 September. The libraries were first built to run on NVIDIA hardware.

According to the FlashMLA README, the attention kernels now power DeepSeek-V4.1 inference on both NVIDIA GPUs and the Ascend 950 NPU. DeepSeek reports 410 TFlops on Ascend 950 prefill, which it puts at 95% of hardware peak, and 360 TFlops on decoding, or 83%. On NVIDIA's B200 it lists up to 1,460 TFlops in prefill. The Ascend build needs CANN 9.2.0 or newer.

MaxForAI said DeepSeek trained its V4 series with many operators written in TileLang, a Python-like language for high-performance kernels, and that every such operator now has an Ascend implementation. The TileLang repository lists Ascend 950 alongside CUDA, AMD ROCm and Apple Metal backends. A post by the account zheanxu, whose affiliation the post does not state, claimed DeepGEMM on Ascend reaches 99.8% of the hardware limit on GEMM and 98% on MegaMoE.

Chinese outlet QbitAI reported further figures from the Ascend side, including DeepEP dispatch at 375 GB/s and combine at 347 GB/s. It said DeepSeek-V4.1-Flash on 32-way expert parallelism reached 5,102 tokens a second per card at a 10-millisecond token interval with a 128K context, in offline mode. These numbers come from DeepSeek and Huawei. No independent benchmark of the Ascend kernels has appeared yet.

Sources 5 sources

  1. Source MaxForAI (X)
  2. Source eliebakouch (X)
  3. Source DeepSeek, FlashMLA README (GitHub)
  4. Source TileLang README (GitHub)
  5. Source QbitAI