Chinese Labs Dominate the Open Model Frontier in 2026

State of Open Models: Summer 2026 Observations

Chinese Labs Dominate the Open Model Frontier in 2026

Hugging Face's biannual analysis of the open model ecosystem from January to August 2026 reveals a dramatic shift: Chinese labs now release the largest and most performant open models, with monthly parameter counts reaching up to 2.78 trillion, while U.S. labs stay under 130B except for a few exceptions. The report highlights that attention (likes) and adoption (downloads) are diverging, with only one model appearing in both top 25 lists. Qwen has become the community's base model with 151,448 derivatives, and small models under 1B still dominate downloads. Open weights are shifting value to APIs, hardware, and ecosystem positions, with Chinese labs licensing their largest models permissively (Apache 2.0 or MIT). The runtime layer, especially llama.cpp, is growing fastest, enabling trillion-parameter models to run locally.

The ceiling moved with llama.cpp. Local inference used to mean an 8B model on a laptop. It now means a trillion-parameter mixture-of-experts spread across a few consumer machines.
  1. cl42

    I recently read a few articles about how the harness is a bigger factor to successful LLM usage and wish they discussed this here.

    I use GLM-5.3, Qwen3.8, Claude (all of 'em), GPT Sol/Luna/Terra across direct API calls + local models where I can (128GB Macbook Pro)... The harness and whether the model or underlying system prompts know how to make the best use of iterative LLM calls makes such a big difference...

    For example: one-off articles on a news topic (e.g., "Update me on the US-Canada relations") yields very similar results across all models... But "run a web search, write a draft perspective from three points of view, and structure data around it" will make everything but Claude + GPT struggle.

  2. AnodicElegy

    The data and graphs are great, but it would have been a much higher quality report if the text and titles were written by a human.

  3. john_rood

    The harness share is the most interesting new data here, but raw request volume can be misleading. A noisy agent with a wide tool loop may generate 10x the Hub calls of a more efficient one. I'd love to see successful outcomes per 1,000 agent-tagged calls, segmented by harness, task class, model, and tool-error rate. Otherwise this tells us which clients are busy, not which ones are effective.

More from this day

2026-09-01