What to Know About Chinese AI Models

Remote Visualization

Recent Chinese AI models are doing well on major benchmarks. The most recent GLM-5.2 model from Z.ai lab performs near the top U.S. closed models in coding and agent tasks. The earlier DeepSeek-R1 release in 2025 similarly shocked the U.S. AI community and challenged the idea of U.S. superiority in the AI race against China. Now, newer releases from Chinese labs suggest that DeepSeek was not a one-shot event but part of a pattern of Chinese labs catching up quickly. This Critical Questions unpacks what recent Chinese model releases mean for capability, cost, open-weight strategy, and U.S. policy.

Q1: Are Chinese AI models capable?

A1: Yes. Chinese models are now close enough to the frontier that they can compete with U.S. models in many real-world tasks. GLM-5.2 is an open-weight model with roughly 750 billion total parameters and a 1 million token context window. It ranks second across other models in the front-end coding benchmark and the top among other open-weight models. Moonshot’s Kimi K2.7 Code model follows OpenAI and Anthropic models closely in the agent benchmarks and software engineering benchmark. The DeepSeek V4-Pro and Qwen3.7-Max models are also highly capable and just behind U.S.-based models. All these recent releases show a larger trend that Chinese AI models are only months, not years, behind U.S. frontier models. The Center for AI Standards and Innovation (CAISI) evaluated open-weight models and found that the gap between the leading U.S. models and DeepSeek V4 Pro is about eight months behind the leading U.S. models.

There are many reasons behind this trend. The most important is how AI models are being trained in Chinese labs. Instead of building the training dataset from scratch, these labs benefit from knowledge distillation. In simple terms, it is a practice of using a stronger model to help train a student model. A stronger model is asked many questions. Its answers, explanations, and judgments are then used as training data for another model. Without fully cloning the stronger model, this process can turn expensive model behavior into cheaper training examples. It helps a new model learn answer style, reasoning patterns, coding habits, and instruction-following behavior without repeating every costly experiment from scratch.

Knowledge distillation is not unique to China. Many U.S. labs also rely on these distillation techniques when training their AI models. DeepSeek itself released distilled R1 models based on Qwen and Llama checkpoints. However, when put into a geopolitical perspective, such practices lead to Chinese labs exploiting frontier U.S. models at scale to train their systems. This is why U.S. companies such as OpenAI and Anthropic accused Chinese labs (DeepSeek, Moonshot, and MiniMax) of stealing their products via 24,000 fraudulent accounts. This is why the White House decided to fight against adversarial distillation as a national security issue and called for information sharing, private-sector coordination, sharing of best practices, and possible accountability measures across U.S. companies.

The second reason for the success of Chinese models is the growing open-weight community and related research. Today, best practices in model architecture, data filtering, post-training, tool use, inference optimization, and evaluation move quickly across the field. Once one lab shows that a technique works, other labs can often imitate it within a short time frame. This allows Chinese labs to learn and implement best practices.

Third, the nature of the current large language model (LLM) race itself narrows the gap. Since the transformer era began in 2017, the main game has been better data, more efficient compute, better post-training, and better inference. There is no permanent secret recipe that gives one lab a qualitative edge forever. Many capabilities can be reached, copied, or imitated in months. China is using the existing research base and open-model infrastructure to shorten the path.

Q2: Why are Chinese models cheaper?

A2: Chinese models are often much cheaper to access than leading U.S. closed models for inference. Chinese labs also claim lower training costs. The best-known example is DeepSeek V3. DeepSeek said the final official training run used 2.788 million H800 GPU hours, which it priced at about $5.6 million. Table 1 shows current public application program interface (API) prices.

Image
Yasir Atalan
Deputy Director and Data Fellow, Futures Lab, Defense and Security Department
Remote Visualization

Several reasons explain the price gap. The first is the pricing mechanism. For closed U.S. models, the model owner controls the API, serving infrastructure, safety layer, compliance system, memory features, tools, connectors, uptime guarantees, and enterprise product. The company largely sets the price. For open-weight models, pricing is more competitive and prices change more frequently. Many providers can host the same model, optimize inference, and compete on price. A model that can be hosted by many providers will face more price pressure than a model that can be accessed only through one owner.

Second, Chinese AI models have lower training costs because they do not have room for error like U.S. companies do. Export controls and weaker access are effective, as they stop Chinese labs from building their own compute infrastructure. This is why Chinese labs work through distillation of U.S. models, open-weight practices, and inference efficiency, reducing the number of expensive experiments needed before release. This suggests that model security tests are more limited compared to U.S. models.

Third, the strategies are different for both sides. Many Chinese labs use open-weight releases to build adoption, prestige, and developer ecosystems. A cheaper model can spread faster. It can become the default for developers who do not need the absolute best frontier model. It can also build global influence even when the company does not control every deployment. By contrast, the U.S. frontier AI is now led by closed-source models, as the United States has the leading edge in compute and therefore does not need to rely on open-weight models.

Q3: Do open-weight models offer a better strategy?

A3: Not necessarily. There are many upsides and downsides. The main benefit is speed of diffusion. Open-weight models can spread through Hugging Face, GitHub, cloud providers, local deployments, and third-party inference platforms. A European startup, a Southeast Asian government agency, or a Latin American developer can use Qwen, DeepSeek, GLM, or Kimi through a third-party host without sending logs directly to the original Chinese lab. This lowers political and operational barriers to adoption. Hugging Face recently announced that Chinese models surpassed U.S. models in both monthly and overall downloads in their platforms, and that Chinese models accounted for 41 percent of their downloads over the previous year. This helps China build prestige and weaken the perception that the United States has an unassailable lead. It also creates de facto technical dependence on Chinese model families, even when the models are hosted outside China. Many users do not care if the model is Chinese or American, as long as they can host them locally.

Open-weight models are also partly a necessity for China. U.S. export controls limit access to the best AI chips, and Chinese labs do not have the same global inference capacity as U.S. hyperscalers. If a Chinese lab had to serve a global user base fully on its own compute, it would face a major capacity problem. Open-weight releases reduce that burden because third-party providers and local users supply much of the serving compute. The model can circulate globally without every request going through China.

But open-weight models cut the other way. Chinese labs do not collect global user data in the same way when their models are used through third-party hosts. If a model is downloaded from a U.S. cloud provider, run through a third-party inference platform, or deployed locally inside a company, the original Chinese developer cannot see the prompts, logs, feedback, tool calls, or product behavior. In other words, monetization is highly difficult with the open-weight strategy.

In the larger context, these diverging strategies will shape the future of AI competition. Models improve when companies can see how users interact with them, where they fail, what tools they connect to, what tasks users repeat, and what outputs users prefer. U.S. firms have an advantage here because they are building integrated AI products around ChatGPT, Claude, Gemini, Microsoft Copilot, enterprise APIs, coding agents, and cloud platforms. Integrated products create feedback loops, enterprise relationships, and product lock-in that open-weight diffusion alone does not create. When Chinese open-weight models spread globally through third-party hosting and local deployment, Chinese labs such as Alibaba, Tencent, Baidu, and ByteDance will have challenges in monetizing the same global interaction data, subscription revenue, and platform lock-in that closed U.S. providers can build.

Q4: What are the implications for U.S. policy?

A4: The United States faces a two-part problem. It needs to both protect its frontier advantage and persuade the rest of the world to build on the American AI stack. The Trump administration’s strategy of exporting the AI stack to the globe is the right direction. The American AI Exports Program aims to promote full-stack AI packages abroad. The Department of Commerce’s call for proposals defines those packages broadly, including AI-optimized hardware, data pipelines and labeling systems, AI models and systems, cybersecurity measures, and sector-specific applications. Such policies are also a strategic response to Chinese open-weight diffusion. If global actors can build on U.S. chips, clouds, models, cybersecurity standards, and applications, the United States will lead in across ecosystem as well.

But this strategy depends on trust. The United States cannot ask other countries to build on its AI stack while giving them the impression that access can be changed suddenly and unilaterally. Recent model-access suspension was not a good sign. On June 12, 2026, the U.S. government suspended Fable and Mythos for foreign access, which led to the full withdrawal of the models by Anthropic. If foreign firms believe U.S. model access can be withdrawn quickly, they will diversify. Some will choose Chinese open-weight models. Some will choose local sovereign models. Some will use multiple providers to avoid dependence on Washington. All this compute without a global user base will not give the United States the advantage it expects.

Prices are also a challenge. Leading U.S. models remain expensive for many developers and governments, whereas Chinese open-weight models offer a cheaper alternative for many enterprises. As the gap between U.S. closed-source models and Chinese open-weight models gets narrower, this price difference will likely lead to a push toward the open-weight models even in the United States. AI companies therefore need to adjust their pricing structure. As for policymakers, U.S. export strategy should focus on decreasing prices and making U.S. models available through more cost-effective options.

Overall, the recent Chinese model releases show that the AI model capability gap is narrow. This is why U.S. AI strategy should be winning global adoption as well as leading through frontier models. The United States still has major advantages, but Chinese models are now capable enough, cheap enough, and open enough to shape the global AI competition. This is why winning the global AI race will depend on the ability to build trust through being a reliable provider.

Yasir Atalan is deputy director and a data fellow with the Futures Lab at the Center for Strategic and International Studies (CSIS) in Washington, D.C.