FR
live

OpenRouter guarantees US residency for requests to Chinese models

Open-weight models accounted for 60% of OpenRouter’s US token consumption in August, most of them Chinese. The marketplace has now made US in-region routing generally available, guaranteeing that requests are decrypted, processed, and served entirely on American soil — or rejected.

A row of identical dark network switch ports, a single amber patch cable plugged into one port in the middle.

August 2026. Open-weight models account for roughly 60% of tokens consumed by US-originating requests on OpenRouter, with Chinese models making up the majority. February 2026. Hugging Face data shows Chinese-developed models took 41% of downloads over twelve months, ahead of the US at 36.5%. September 14, 2026. OpenRouter moves US in-region routing into general availability, promising to decrypt, process, and serve every request entirely inside the United States — or reject it. Why it matters: data sovereignty is no longer a question of which model you use but of how it is routed, and a single URL is now enough to guarantee it.

The open-weight paradox

The open-weight pitch is well rehearsed: a company downloads the weights, customizes them, runs them on infrastructure of its choosing, and keeps far greater control over where its data is processed — often at a lower cost than proprietary models. Open-weight models are now thought to trail the leading frontier systems by only four to five months, and NVIDIA, the world’s most valuable company, is betting heavily on that future. In early September it agreed to acquire Hugging Face, the “GitHub of AI” hosting more than three million models, for $12.9 billion, pledging to keep the platform open.

That power cuts both ways. OpenAI president Greg Brockman recently warned that increasingly capable open-weight models — pointing specifically at China’s GLM-5.3 — could “significantly accelerate the threat landscape” as models with advanced cyber capabilities become freely downloadable and modifiable.

But for a business accessing those models through a third-party service, a more operational question looms: where does its own data go when it uses a model, especially one developed in China.

A middleman that controls the geography

OpenRouter was launched in early 2023 by Alex Atallah, the former CTO of OpenSea. It is an interface to the crowded AI model market: a developer switches between hundreds of models from dozens of providers through a single API. Payments giant Stripe recently announced plans to acquire it in a reported $8 billion deal, while Cursor, Ramp, and Meta are building their own routers.

The appeal of model routers comes down to economics. A developer who once hard-coded everything to the same model can now choose request by request, sending easier jobs to cheaper models while reserving pricier frontier systems for the work that actually needs them.

That intermediary role is exactly what makes the new residency controls possible. OpenRouter already decides which provider serves each request, so it can now restrict that choice to provider endpoints operating in the United States.

What in-region routing changes

With the standard global endpoint, a request can be served by an eligible provider in any region — so using a model from a US company does not even guarantee the request is processed in the US. With us.openrouter.ai, the request is decrypted on OpenRouter’s infrastructure inside the US and the provider pool is filtered to endpoints approved as operating there.

If no compliant US provider can serve the requested model, OpenRouter returns a 404 error. Companies can also enforce the restriction through OpenRouter’s Guardrails at the workspace, team, or API-key level, and tools that would send prompt data outside the US are disabled on the regional endpoint.

The feature had already existed quietly: OpenRouter’s documentation has said since early August that US in-region routing was available to enterprise customers on request. It joins European in-region routing, available since October 2025.

The European precedent shows how this generalizes. Since October 2025, OpenRouter has offered the same guarantee for the EU, and the mechanism is identical: a regional endpoint, a filtered provider pool, and a hard failure when no compliant provider exists. The point worth stressing is that residency is enforced at the routing layer, not promised by the model vendor. A model from a US lab can be served from a European data center, and a Chinese model from a US one — the guarantee follows the provider endpoint, not the model’s origin. That is why the eligible-model list matters more than any single vendor’s marketing claim.

Chinese models, still Chinese — but served from the US

Cailee Moberg, on OpenRouter’s product team, sums up the customer dilemma: “Models from Chinese labs are still most of the [open-weight model] volume, and procurement approval for those models can be difficult.” OpenRouter’s answer is to let teams capture the price and performance gains of Chinese open-weight models without giving up data residency.

DeepSeek V4 Pro, Kimi K3, and GLM 5.2 are all available through US in-region routing because Baseten, Fireworks, and Azure serve them from US data centers. A company could already keep these models inside the US by self-hosting them or going through a US provider directly; OpenRouter’s routing gives its own customers that guarantee without managing those deployments themselves.

The tailwind is real. In its State of AI in the Enterprise 2026 report, Deloitte found that 77% of companies “now factor country of origin into their vendor selection,” while nearly 60% build their AI stacks “primarily with local vendors.”

Sovereignty, regulation, and compliance

In-region routing does not appear out of nowhere: it answers a hardening regulatory framework. The EU’s GDPR already imposes safeguards on transfers of personal data outside the Union, and the US CLOUD Act lets authorities demand access to data held by US providers, wherever it is stored. For a European company wanting to use a Chinese model through a US provider, the chain of accountability quickly becomes unreadable.

Deloitte’s number confirms it: 77% of companies factor country of origin into vendor selection, and nearly 60% build their AI stacks “primarily with local vendors.” Sovereignty is no longer a theoretical debate — it is a procurement constraint.

OpenRouter’s guarantee covers one specific link: where the prompt is decrypted and processed. It does not cover the model’s training, nor the lab that designed it, nor the agreements between the US provider and that lab. It is a guarantee of processing, not of provenance. For a DPO, it is one item for the compliance file, not a waiver.

The setup is deliberately small: point the client at the regional endpoint, enforce the constraint in Guardrails, and the residency guarantee is in force. The main cost is flexibility — a model without a compliant US provider simply becomes unavailable, which is the point of the feature.

Verdict

OpenRouter’s US in-region routing does not change where the models come from — it changes which copy your requests are routed to and where your prompts are handled along the way.

If you are under a data residency requirement and want to use Chinese open-weight models, switch your calls to us.openrouter.ai and enforce the constraint through Guardrails at the API-key level. If you have no regulatory constraint, the global endpoint remains more flexible and possibly cheaper — but first measure where your requests actually go. If you handle highly sensitive data, a routing guarantee is not enough: self-hosting or a direct US provider remains the only option that removes the intermediary entirely.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

Amazon Kinesis Data Streams now writes Apache Iceberg tables directly

On August 31, 2026, AWS launched streaming tables for Kinesis Data Streams, turning a stream into a queryable Apache Iceberg table with no pipeline to operate. The vendor claims up to 50% delivery savings and 30% query savings — at the cost of delegating compaction.

AWS open-sources Pizza Bot, an email inbox for background AI agents

On September 10, AWS open-sourced Pizza Bot, an app that replaces chat with an email-style inbox for tracking AI agents working in the background. For any team running autonomous agents in the cloud, the thesis fits in one sentence: the interface must assume no one is watching.

← Back to the feed

Type at least two characters.

navigate open esc dismiss