Performance evidence · Reviewed 2026-09-22
What has actually been measured?
No verified latest-versus-latest independent comparison
Follow-up through September 22 did not verify a matched independent comparison of current Chinese accelerators against Rubin, GB300, MI355X, TPU v7, or OpenAI Jalapeño. The review checked MLPerf v6.0 repository heads and submitter listings plus InferenceX's published coverage, not every result artifact. A Pengcheng Laboratory preprint adds measured Ascend 910C scientific-computing results against A800/H800 controls, but it is not a model-training or inference comparison and does not use current frontier foreign hardware. Absence of a verified public result is not evidence of incapability.
Published hardware specifications
Dense BF16 peak per marketed accelerator and HBM bandwidth. These are specifications, not model training or inference benchmarks. Package design and power differ.
Scroll sideways to compare all columns →
DeepSeek-R1 on Ascend 910C and H100
2025-06 · vendor study · Historical comparison
Scroll sideways to compare all columns →
Vendor/partner measurements versus published NVIDIA baselines. INT8 versus FP8, different decode batches, and speculative-decoding assumptions prevent a universal ratio. Perfect expert-balancing results are projections.
Huawei / SiliconFlow · Serving LLMs on Huawei CloudMatrix384 (source)Benchmark evidence register
2026-07-22measured
Ascend 910 scientific-computing study
Pengcheng Laboratory measured Ascend 910A/B/C across scientific kernels and five applications. In one matched quantum-simulation case, a tuned 910C implementation completed a 30-qubit, 30-layer circuit in 11.4 seconds versus 14.3 seconds on A800 with cuQuantum. The study also reports weaker vector and memory-bound behavior, relies on workload-specific optimization, compares against export-limited A800/H800 rather than current frontier chips, and does not establish AI model-training or inference parity.
Pengcheng Laboratory · Ascend to Science: Exploration of AI Chips for Scientific Computing (source)2026-09-07vendor
Biren GLM-5.3-Flash functional validation
Vendor reports functional and accuracy checks on Bili 166M with SGLang and BIRENSUPA. Compatibility evidence only; no matched foreign-hardware throughput, latency or power comparison.
Biren · Biren reports GLM-5.3-Flash inference validation on Bili 166M (source)2026standardized
MLPerf Training / Inference v6.0
Standardized submitter-run results; no matched latest Chinese-versus-foreign comparison established in this baseline.
MLCommons · MLPerf Inference v6.0 submissions (source)MLCommons · MLPerf Training v6.0 submissions (source)Continuousmeasured
InferenceX
Reproducible independent inference runs. The current dashboard names Rubin, GB300, MI355X, TPU v7 and OpenAI Jalapeño systems; equivalent current Chinese results were not verified.
SemiAnalysis · Continuous inference benchmark dashboard (source)2026-07measured
Ascend operational field study
A small independent preprint documents deployment friction on older Ascend 910 devices. No matched NVIDIA control.
Zheng Yu, CUHK Shenzhen · Independent Ascend deployment field study (source)2026-04vendor
HiFloat4 pretraining
Demonstrates training progress on dense and MoE models; does not measure frontier time-to-quality against foreign hardware.
Huawei · HiFloat4 language-model pretraining study (source)How to read these results
Independent means the evaluator ran the workload independently of the vendor. MLPerf provides standardized rules but its submissions are generally run by submitters. Vendor and partner studies remain attributed to their authors.
A useful comparison aligns model, quality, precision, context, batch, latency target, accelerator count, software and power. Training needs time-to-quality and reliability; inference needs both throughput and latency. Missing evidence does not establish incapability.