Performance evidence · Reviewed 2026-09-22

What has actually been measured?

No verified latest-versus-latest independent comparison

Follow-up through September 22 did not verify a matched independent comparison of current Chinese accelerators against Rubin, GB300, MI355X, TPU v7, or OpenAI Jalapeño. The review checked MLPerf v6.0 repository heads and submitter listings plus InferenceX's published coverage, not every result artifact. A Pengcheng Laboratory preprint adds measured Ascend 910C scientific-computing results against A800/H800 controls, but it is not a model-training or inference comparison and does not use current frontier foreign hardware. Absence of a verified public result is not evidence of incapability.

Published hardware specifications

Dense BF16 peak per marketed accelerator and HBM bandwidth. These are specifications, not model training or inference benchmarks. Package design and power differ.

Scroll sideways to compare all columns →

AcceleratorDense BF16
(PFLOP/s)
HBM bandwidth
(TB/s)
Basis & limitations
MI355XAMD · Outside China2.58Dense BF16 peak; high-power accelerator module.AMD · MI355X specifications (source)
TPU v7 / IronwoodGoogle · Outside China2.3077.38Per marketed chip, containing two compute chiplets.Google Cloud · TPU7x Ironwood specifications (source)
Ascend 910CHuawei · China0.7523.2Dual-die accelerator package; specifications from the 2025 paper.Huawei / SiliconFlow · Serving LLMs on Huawei CloudMatrix384 (source)
Atlas 350 / 950PRHuawei · China0.4251.4Inference/prefill-oriented PCIe card, up to 600 W.Huawei · Atlas 350 accelerator specifications (source)
Blackwell Ultra / GB300NVIDIA · Outside China2.58Per GPU, derived from 72-GPU system totals; sparsity removed from BF16 peak.NVIDIA · GB300 NVL72 specifications (source)

DeepSeek-R1 on Ascend 910C and H100

2025-06 · vendor study · Historical comparison

Scroll sideways to compare all columns →

MetricAscend 910CH100 baseline
Prefill tokens/s per accelerator5,6556,288
Decode tokens/s per accelerator1,9432,172
Decode time per output token49.4 ms55.6 ms

Vendor/partner measurements versus published NVIDIA baselines. INT8 versus FP8, different decode batches, and speculative-decoding assumptions prevent a universal ratio. Perfect expert-balancing results are projections.

Huawei / SiliconFlow · Serving LLMs on Huawei CloudMatrix384 (source)

Benchmark evidence register

2026-07-22measured

Ascend 910 scientific-computing study

Pengcheng Laboratory measured Ascend 910A/B/C across scientific kernels and five applications. In one matched quantum-simulation case, a tuned 910C implementation completed a 30-qubit, 30-layer circuit in 11.4 seconds versus 14.3 seconds on A800 with cuQuantum. The study also reports weaker vector and memory-bound behavior, relies on workload-specific optimization, compares against export-limited A800/H800 rather than current frontier chips, and does not establish AI model-training or inference parity.

Pengcheng Laboratory · Ascend to Science: Exploration of AI Chips for Scientific Computing (source)

How to read these results

Independent means the evaluator ran the workload independently of the vendor. MLPerf provides standardized rules but its submissions are generally run by submitters. Vendor and partner studies remain attributed to their authors.

A useful comparison aligns model, quality, precision, context, batch, latency target, accelerator count, software and power. Training needs time-to-quality and reliability; inference needs both throughput and latency. Missing evidence does not establish incapability.