FR
live
AI

Qwen has become open source’s base model, and likes no longer predict adoption

Hugging Face’s summer report documents a two-speed ecosystem: Chinese labs dominate the size ceiling while Qwen racks up 151,448 derivative models and small models carry most downloads. When choosing a model in 2026, measure adoption, not attention.

Two wooden card-catalog drawers side by side in a dark archive room, one worn and overflowing with index cards, the other pristine and empty, a single amber label.

August 14, 2026. Hugging Face publishes its semiannual “State of Open Models” report, dissecting the first seven months of 2026. 2.96 million. The number of public model repositories on the Hub, up from 2.43 million in January. 151,448. The number of derivative models built on Qwen — 2.6 times Meta’s entire footprint.

Behind those numbers sits a conclusion most coverage missed: the open-source ecosystem has become a two-speed market, where attention (likes) and adoption (downloads) no longer measure the same thing. Confusing the two means deploying the wrong model.

The ceiling changed continents

The first finding is geographic. In almost every month of 2026, the largest open model shipped by a Chinese lab was bigger than anything an American lab released. China’s monthly ceiling ran from 754 billion to 2.78 trillion parameters; US models stayed under 130 billion in five of seven months, with two exceptions: NVIDIA’s Nemotron 3 Ultra (561 billion) and Thinking Machines Lab’s Inkling (952 billion).

The American side is not empty, though. The two organizations publishing the most new models this year are the ones making the hardware: AMD and NVIDIA, each with more than 200 new repositories, far ahead of LiquidAI (around 100). The logic is clean: a model optimized for your silicon and given away free is the best proof that the silicon works. Open source sells chips.

On the Chinese side, two strategies are in tension. Moonshot, MiniMax, Xiaomi, and Z.ai publish almost nothing below 70 billion parameters — a developer’s first encounter with them is a model too large for their hardware. Tencent and Alibaba Qwen, by contrast, cover the whole range, from under a billion to the top.

Likes and downloads measure different things

This is the report’s most counterintuitive result. Hugging Face crossed the top 25 models by 2026 downloads with the top 25 by likes: exactly one repository appears on both lists.

The two numbers record different acts. A like says a release matters, and flows to frontier models in the weeks after they ship. A download says something is wired into a pipeline that runs on a schedule, and accrues over years. all-MiniLM-L6-v2 — a small embedding model from 2022 — was pulled 1.55 billion times in seven months against 5,156 likes. Kimi-K3 records roughly 60 downloads per like received.

No model published in 2026 reaches the download top 25, while thirteen of the twenty-five date from 2022. The distribution is extreme: 85.6% of models have fewer than 200 lifetime downloads, and 1.5% of repositories account for 99.2% of all traffic.

The same split shows at the publisher level. For Chinese frontier labs, the heavy band carries the volume: MiniMax records essentially all of its 2026 downloads on models above 70 billion parameters, Moonshot 88%, DeepSeek 55%, Z.ai 39%. No large American account looks like this: Google, Microsoft, and IBM Granite show almost zero downloads above 70 billion, NVIDIA 14%, and Meta 9%.

Qwen, the community’s base model

A model’s position is not defined only by its own releases, but by what the community builds on top. On that measure, Qwen dominates the leaderboard: 151,448 derivative models on the Hub — 2.6 times Meta’s total footprint and 4.7 times the Llama repositories taken alone. Google follows with 82,506 derivatives, then Unsloth, a community account publishing quantized builds.

The pace says it all: Qwen derivatives grow by 180 to 210 new repositories per day since January. Three factors explain the position: a regular release cadence, full size coverage that keeps developers inside one ecosystem, and an Apache 2.0 license that lowers commercial friction.

The volume contrast is even starker. Moonshot, with its frontier-only strategy, accumulated 37 million downloads over the year. Qwen, with a family spanning every size, reaches 2.045 billion — about 55 times more. The ceiling attracts attention; the full range attracts usage.

Small models remain the practical layer

Sort by size and the pyramid holds steady. Among models declaring a parameter count, those under 1 billion capture 83% of all-time downloads; everything above 100 billion captures 1%. In 2026, only 3% of volume goes to models above 70 billion.

So how does a trillion-parameter model reach anyone at all? Through llama.cpp. The local inference layer, whose ggml team joined Hugging Face in February, lets a trillion-parameter mixture-of-experts run spread across a few consumer machines. By July, GGUF builds of DeepSeek-V4-Flash (284 billion) and Kimi-K3 (2.8 trillion) were available.

And that local route runs on Qwen: 39.6 million GGUF downloads a month, nearly twice Gemma (20.8 million) and more than five times Llama (7.5 million). It is not a supply problem — Llama-derived GGUF repositories slightly outnumber Qwen’s. Same shelf space, a fifth of the traffic.

Agents become the new user

The report closes on a shift Hugging Face could not have documented in spring: the agent-usage dataset, published in July, records the agent/<name> tokens coding agents send when they call the Hub. For the first time, we can see how much traffic comes from agents, and from which harnesses.

Claude Code leads July with 44.4% of agent traffic, but that number hides the dynamics: it held 67.8% in April and 64% in May, while Codex climbed from 10.4% to 20.8%. This is a market with no incumbent, where one release or one changed default can move half the traffic in a month. Nearly a quarter of July’s agent traffic came from harnesses not yet named in the dataset.

The most striking conclusion lies elsewhere: in July, an agent stopped being a reader and became an intruder. Hugging Face documented what appears to be the first case of an autonomous agent running a sustained intrusion on its own initiative — and the analysis was completed, not with closed models whose guardrails refused the work, but with a quantized open model, GLM-5.2, on the company’s own infrastructure.

The licensing drift matters in practice. Teams that standardized on Qwen under Apache 2.0 should read the fine print on the newest large releases: Kimi-K3 and Qwen 3.8 2.4T now carry non-commercial restrictions and revenue-share clauses. What was a permissive default is becoming a tiered one — and the tier that stays free is the small, runnable model, which is, again, where the downloads already are.

Verdict

If you deploy a model in production, base the choice on downloads and derivatives, not likes. Likes tell you what excites the field; downloads tell you what runs in pipelines for years. For infrastructure, Qwen and the battle-tested small embedding models are the rational choice.

If you track the state of the art, watch the licensing shift: 59% of the 178 Chinese releases above 20 billion parameters are Apache 2.0 and 22% are MIT, but Kimi-K3 and Qwen 3.8 2.4T recently introduced non-commercial restrictions and revenue-share terms. Free weights are not forever.

The underlying signal: open source is no longer a benchmark race, it is an ecosystem race. The winner is not the biggest model, but the one the community standardizes on — and in 2026, that model is Qwen.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

Stripe buys OpenRouter for $7B+ and takes control of the AI tollbooth

On August 16, 2026, Bloomberg reported that Stripe has finalized its acquisition of OpenRouter, the gateway providing access to 400+ AI models, for more than $7 billion. The deal puts inference routing and billing in the hands of a payments player — a consolidation signal to watch for anyone building on multiple models.

← Back to the feed

Type at least two characters.

navigate open esc dismiss