ChinAI #305: Computing Power Shifts in the AI Inference Era
Greetings from a world where…
you are still worthy even if you don’t produce
…As always, the searchable archive of all past issues is here. Please please subscribe here to support ChinAI under a Guardian/Wikipedia-style tipping model (everyone gets the same content but those who can pay support access for all AND compensation for awesome ChinAI contributors).
Feature Translation: DeepSeek has sparked a crazy rush for Nvidia H20s, but the AI inference explosion is not just about hoarding chips
Context: In recent months, the NVIDIA H20 chip — once neglected by Chinese tech companies — has become a hot commodity. Tencent recently ordered 100,000 to 200,000 of these chips. Why? This week’s Qbit AI report (link to original Chinese) unpacks the shifts in computing power in the AI inference era. I’ve been impressed by Qbit’s coverage on this topic.
A few issues ago, ChinAI featured Lily Li’s translation (ChinAI #299) of another report that asks: If we use daily average consumption of 1 billion tokens as a threshold for “beginner-level” of large language model adoption, what can we learn about the diffusion of AI in China?
Key Takeaways: DeepSeek has reconfigured the basis of computing power for AI development, shifting the “training-based” paradigm to an “inference-based” one.
Compared with training models with more and more parameters, running and deploying these models (the inference stage) requires fewer computing resources, in terms of 1) hardware thresholds and 2) ultra-large-scale cluster construction.
Let’s start with hardware thresholds because that’s where the H20 comes in. The article states, “Although the performance of the H20 is only 1/10 of that of the H100, it is more than enough for inference, has enough memory bandwidth, is suitable for running large-scale parameter models, and is even cheaper.”
This article largely draws on an interview with Yao Xin, CEO of PPIO [派欧云] (a Chinese edge cloud computing services company ), so keep in mind his biases. From the article: “Yao Xin judged that in the future, NVIDIA's monopoly will also change. In the inference era, inference chips will flourish. For example, according to the test results of DeepSeek researchers, the performance of (Huawei’s) Ascend 910C in inference tasks can reach 60% of that of H100.”
Note, however, last week’s ChinAI issue, which emphasizes that some of these comparisons of Chinese chips to NVIDIA chips don’t specify whether it is the full-parameter or distilled version of DeepSeek.
In the AI inference era, computing demands are also changing when it comes to cluster construction. In other words, this is not a race to trillion-dollar clusters.
Edge cloud computing providers like PPIO run AI services through distributed networks, leveraging 3900+ computing power nodes from across China. Qbit reports that, during Chinese New Year, “PPIO achieved 99.9% availability of its To-Business DeepSeek services…At present, the average daily token consumption of the PPIO platform has exceeded 130 billion, which is comparable to the average daily token consumption of the ‘Six Little Dragons’ (China’s top frontier AI startups).”

PPIO has also provided large-scale AI inference services for Baichuan Intelligence, but this is just one pathway to AI inference. In contrast, Alibaba distributes its AI models through its own internet data centers.
In a way, the architecture of edge cloud computing is similar to DeepSeek’s core algorithmic innovations. Yao reflects, “The cross-node expert parallel system proposed by DeepSeek has already reflected the idea of distribution to a certain extent. It concentrates uncommonly used expert models on one server, and allocates more computing power to commonly used expert models. This forms a balance in scheduling.”
Full Translation: DeepSeek has sparked a crazy rush for Nvidia H20s, but the AI inference explosion is not just about hoarding chips
ChinAI Links (Four to Forward)
Should-read: What Does the Public Think About AI?
This GovAI report synthesizes public attitudes toward AI and introduces “the AI Survey Hub for Attitudes and Research Exchange (AI SHARE) database, which aggregates survey data from over 200 studies conducted between 2014 and 2023.” Authors: Noemi Dreksler, Harry Law, Chloe Ahn, Daniel S. Schiff, Kaylyn Jackson Schiff, Zachary Peskowitz
Should-read: How Might the United States Engage with China on AI Security Without Diffusing Technology?
In a RAND commentary, Karson Elmgren derives some practical implications on U.S.-China AI safety and security cooperation. He cites my EJIR article on the historical lessons from U.S. cooperation with the Soviet Union during the Cold War on permissive action links (PALs). He writes, “United States may again wish to keep its competitors safer to assure its own safety.”
Should-read: The role of human-capital in artificial intelligence adoption
This Economics Letters article concludes: “Using cross-country-cross-industry-level information on AI-adoption from Eurostat, we find that higher levels of human-capital are positively associated with higher AI adoption 10 years later. This effect is driven by the most human-capital-intense firms and explains up to half of the observed differences in AI-adoption across European countries and industries.”
Should-read: Large AI models are cultural and social technologies
Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans in Science: “We now have a technology that does for written and pictured culture what large-scale markets do for the economy, what large-scale bureaucracy does for society, and perhaps even comparable with what print once did for language.”
Thank you for reading and engaging.
These are Jeff Ding's (sometimes) weekly translations of Chinese-language musings on AI and related topics. Jeff is an Assistant Professor of Political Science at George Washington University.
Check out the archive of all past issues here & please subscribe here to support ChinAI under a Guardian/Wikipedia-style tipping model (everyone gets the same content but those who can pay for a subscription will support access for all).
Also! Listen to narrations of the ChinAI Newsletter in podcast format here.