Introduction
The latest weekly AI-model usage figures show Chinese models continuing to attract heavy traffic on OpenRouter.
For the week of August 3–9, 2026, media calculations based on OpenRouter data put total token usage across the tracked models at 69 trillion tokens, up 21.48% week over week.
Models developed by Chinese companies accounted for 34.25 trillion tokens, a 21.76% weekly increase. Models classified as U.S.-developed accounted for 9.17 trillion tokens during the same period.
Under that country-level calculation, Chinese models have now exceeded U.S. models for 15 consecutive weeks.
The strongest individual performer was DeepSeek-V4-Flash-0731, which reached 8.83 trillion tokens in weekly usage and rose 570% from the previous week. Tencent's Hy3 followed with 8.05 trillion tokens, up 67%.
Data-scope note: These figures describe usage observed through OpenRouter and country-level calculations derived from its model rankings. They should not be interpreted as the total worldwide usage of every AI model across first-party apps, direct APIs, private deployments, enterprise clouds, and local inference.
Chinese Models Accounted for 34.25 Trillion Weekly Tokens
The original report highlights the scale of Chinese-model usage in the latest OpenRouter snapshot.
The main figures were:
| Metric | August 3–9, 2026 | Week-over-Week Change |
|---|---|---|
| Total tracked model usage | 69T tokens | +21.48% |
| Chinese-model usage | 34.25T tokens | +21.76% |
| U.S.-model usage | 9.17T tokens | +109.36% |
The Chinese total represented almost half of the 69 trillion tokens cited for the week.
This was also the 15th consecutive week in which the Chinese-model total exceeded the U.S.-model total in the calculation reported by Chinese media using OpenRouter data.
That run is notable because OpenRouter is a multi-model API routing platform rather than a single model provider. Developers use it to access models from many companies through a unified interface, which makes its usage rankings a useful signal for third-party developer demand.
OpenRouter says its rankings are based on real usage from millions of people accessing models through its platform.
However, OpenRouter is only one distribution channel. A model may have very large first-party traffic that never passes through OpenRouter. Tencent, DeepSeek, OpenAI, Anthropic, Google, and other providers also serve users through their own APIs and products, while many open-weight models run privately or locally.
For that reason, the data is best read as a strong indicator of OpenRouter ecosystem usage, not as a complete census of global AI consumption.
DeepSeek-V4-Flash-0731 Jumped 570% and Took First Place
The standout model of the week was DeepSeek-V4-Flash-0731.
Its OpenRouter weekly usage reached 8.83 trillion tokens, an increase of 570% from the previous week, putting it at the top of the cited ranking.
The surge followed DeepSeek's July 31 update to the official V4-Flash API.
DeepSeek's official changelog says the new DeepSeek-V4-Flash-0731 keeps the same model architecture and size as the V4-Flash preview and was updated through additional post-training rather than a new base architecture.
The company also says the release substantially improves agent capabilities and adds native support for the Responses API, including targeted compatibility with coding-agent workflows.
DeepSeek-V4-Flash-0731 at a Glance
Official DeepSeek documentation currently lists the following characteristics:
| Specification | DeepSeek-V4-Flash-0731 |
|---|---|
| API model name | deepseek-v4-flash |
| Context length | 1M tokens |
| Maximum output | Up to 384K tokens |
| Thinking mode | Supported |
| Non-thinking mode | Supported |
| Tool calls | Supported |
| Responses API | Supported |
| Anthropic-compatible API | Supported |
DeepSeek describes V4-Flash as its faster and more economical V4 option. The V4 family originally launched with a 284B-total / 13B-active Flash configuration.
The July 31 revision focused particularly on agent workloads. DeepSeek's published benchmark table reports gains across tasks such as Terminal Bench, DeepSWE, CyberGym, Toolathlon, and internal full-stack coding evaluations.
Those technical improvements provide useful context for the usage jump. Coding agents and tool-using systems can consume far more tokens per task than ordinary short-form chat because one job may involve repeated planning, tool calls, file inspection, code generation, execution, and revision.
A sharp rise in token usage therefore does not necessarily mean the same percentage increase in unique users. It may reflect a mixture of new adoption, longer contexts, heavier agent workflows, greater availability, provider pricing, and repeated machine-to-machine calls.
Tencent Hy3 Reached 8.05 Trillion Tokens
Tencent's Hy3 ranked second in the source data with 8.05 trillion tokens for the week, up 67% week over week.
Hy3 had already shown strong OpenRouter momentum shortly after launch.
Tencent says the model recorded more than 68 times as many API calls as Hy2 within one week of its release and reached first place on OpenRouter's usage leaderboard during that early period.
The model was officially released on July 6 and expanded to broader international availability on August 5.
Hy3 at a Glance
Tencent describes Hy3 as a hybrid fast-and-slow-thinking Mixture-of-Experts model with:
| Specification | Tencent Hy3 |
|---|---|
| Total parameters | 295B |
| Active parameters | 21B |
| Context length | Up to 256K |
| Architecture | Mixture of Experts |
| Reasoning | Configurable fast/slow reasoning |
| License | Apache 2.0 |
Tencent positions Hy3 for coding, document processing, financial modelling, front-end design, game production, and agent workflows.
The company has also integrated the model into products such as WorkBuddy, Tencent Design Miora, and Tencent Cloud TokenHub, while making it available through third-party developer platforms and open model communities.
That wider distribution may help explain why Hy3 remains near the top of third-party usage rankings rather than depending on one first-party application.
Why Token Usage Is a Useful Metric — and Why It Can Be Misleading
Token volume is a useful way to measure how much model inference is happening, but it is not the same as measuring users, revenue, intelligence, or market share.
One developer running a long coding agent can consume more tokens than hundreds of people asking short questions.
A model with a one-million-token context window may process enormous inputs in a single request. Another model may achieve the same task with fewer tokens. Different tokenizers also segment text differently, so one token count is not always directly equivalent to another model's count.
Several factors can push token usage upward:
- More users
- More API integrations
- Longer contexts
- Coding-agent loops
- Tool calling
- Lower prices
- Free or discounted access
- Better provider availability
- High-volume automated workloads
- Repeated retries or inefficient prompting
For that reason, token usage should be interpreted alongside other signals such as request count, unique users, revenue, latency, cost per task, benchmark performance, and retention.
The original report's main point still holds within its data source: Chinese models are receiving unusually heavy usage on OpenRouter, and the trend has persisted for many weeks.
What the 15-Week Run Suggests
A one-week spike can be caused by a launch, a promotion, or temporary pricing.
Build a showcase site and grow leads in minutes
Describe your idea once, and We0 AI can generate a showcase site, pages, and CMS, then help you attract customers and traffic after launch.
One complete project generation for free registration
Best for trying one complete generation flow and seeing a first project draft quickly.
Fifteen consecutive weeks is a more durable signal.
The sustained position suggests that Chinese models are no longer being used only because they are inexpensive alternatives. Developers are repeatedly selecting them for production and experimental workloads across coding, agents, reasoning, and general API use.
Three factors appear especially relevant.
- Strong Price-to-Performance Competition
DeepSeek and Tencent have both emphasized cost efficiency while improving reasoning and agent capabilities.
For API developers, the practical comparison is often not simply which model has the highest benchmark score. It is which model can complete a task reliably at an acceptable total cost.
A model that is slightly weaker on one benchmark may still win substantial traffic if it is faster, cheaper, easier to integrate, or good enough for high-volume workloads.
- Open and Multi-Platform Distribution
Hy3 is released under Apache 2.0, and DeepSeek has also built a strong open-model ecosystem around its releases.
Open and broadly distributed models can appear across many gateways, clouds, local runtimes, coding tools, and agent frameworks.
That distribution reduces dependence on a single product interface and makes adoption easier for developers who already have multi-model infrastructure.
- Agent Workloads Are Becoming a Major Source of Tokens
Both DeepSeek-V4-Flash and Hy3 are marketed heavily around agent capabilities.
Agent workloads are naturally token intensive because the model may repeatedly inspect context, make plans, call tools, read results, and continue until the job is complete.
As coding assistants and autonomous workflow tools become more common, weekly token rankings may increasingly reflect which models are popular inside software systems, not only which chatbots are popular with consumers.
A More Accurate Way to Read the Headline
The source headline says China's AI ecosystem is having a major moment, with model usage leading globally.
The underlying data supports a narrower and more defensible statement:
Chinese-developed models led the cited OpenRouter weekly token-usage calculation for the 15th consecutive week.
That is still significant.
OpenRouter is a major multi-model distribution layer, and high usage there shows that developers are actively choosing these models when they have alternatives from many vendors.
But it would be inaccurate to use the 34.25 trillion-token figure as proof that Chinese models generated more tokens than U.S. models across the entire global AI market.
First-party traffic from ChatGPT, Gemini, Claude, DeepSeek, Tencent products, enterprise deployments, private inference, and local open-weight installations is not comprehensively represented in the OpenRouter total.
The strongest conclusion is therefore about developer-platform momentum, not absolute worldwide AI consumption.
FAQ
What was the total AI model usage reported for August 3–9, 2026?
The cited calculation based on OpenRouter data reported 69 trillion tokens for the week, up 21.48% from the previous week. This is platform-derived usage data and should not be treated as the total token consumption of the entire global AI industry.
How many tokens were attributed to Chinese AI models?
Chinese-developed models were calculated at 34.25 trillion tokens for the week, up 21.76%. The same media calculation put U.S.-developed models at 9.17 trillion tokens.
How long have Chinese models been ahead of U.S. models in this ranking?
The August 3–9 period marked the 15th consecutive week in which the reported Chinese-model total exceeded the U.S.-model total. The comparison is based on OpenRouter usage data and the media's country classification of listed models.
Which model ranked first?
DeepSeek-V4-Flash-0731 ranked first in the cited weekly snapshot with 8.83 trillion tokens. Its usage increased 570% week over week.
What is DeepSeek-V4-Flash-0731?
It is the July 31 post-trained update to DeepSeek's V4-Flash API model. DeepSeek says it keeps the same architecture and size as the preview version while significantly improving agent capabilities.
How much usage did Tencent Hy3 record?
The source reported 8.05 trillion tokens for Hy3 during the week, an increase of 67%. Tencent had already reported strong OpenRouter adoption shortly after the model's July release.
Does OpenRouter data represent all AI usage worldwide?
No. OpenRouter measures activity routed through its own multi-model platform. Direct provider APIs, first-party consumer apps, enterprise clouds, private deployments, and local inference may not be included.
Is token volume the same as model popularity?
Not exactly. Token volume can reflect popularity, but it is also affected by context length, agent loops, pricing, workload type, and token efficiency. Unique users, request counts, revenue, and retention provide different views of adoption.
Related Tools
- OpenRouter: A unified API platform for accessing and routing requests across hundreds of AI models and providers.
- OpenRouter Rankings: Live model rankings and usage views based on activity across the OpenRouter platform.
- DeepSeek API: DeepSeek's official developer platform for accessing V4-Flash and other API models.
- DeepSeek API Documentation: Official model, pricing, tool-use, context, and integration documentation.
- Tencent Hy3: Tencent's official Hy3 model site and access information.
- Tencent Cloud TokenHub: Tencent's multi-model API and model-orchestration platform.
- WorkBuddy: Tencent's AI agent workspace with Hy3 integration.
Related Links
- Original AIBase Report: The Chinese source article used as the editorial starting point.
- OpenRouter AI Model Rankings: The primary platform page for current model usage and ranking data.
- OpenRouter Models: Model availability, descriptions, providers, context windows, and pricing on OpenRouter.
- DeepSeek V4-Flash July 31 Update: DeepSeek's official changelog for the V4-Flash-0731 release and agent benchmark updates.
- DeepSeek Models and Pricing: Official V4-Flash model identifiers, context limits, supported APIs, and current pricing information.
- Tencent Hy3 Global Availability: Tencent's official description of Hy3 adoption, architecture, integrations, and OpenRouter momentum.
- Tencent Hy3 Official Release: Tencent's July release announcement with Hy3 architecture and capability information.
Summary
OpenRouter usage data for August 3–9, 2026 showed another strong week for Chinese-developed AI models. The cited calculation put them at 34.25 trillion tokens, marking a 15th consecutive week ahead of U.S.-developed models within that dataset.
DeepSeek-V4-Flash-0731 led the individual model ranking with 8.83 trillion tokens after a 570% weekly surge, while Tencent Hy3 followed at 8.05 trillion tokens with 67% growth.
The numbers are a meaningful signal of developer adoption, especially for agent and coding workloads, but they are not a complete measure of worldwide AI usage. OpenRouter does not capture every first-party API, consumer product, private deployment, or local model run.
The clearest takeaway is that Chinese models have established sustained momentum in the OpenRouter ecosystem, with DeepSeek and Tencent currently driving a large share of that usage.



