Two numbers about open-weight models were published within weeks of each other this year, and they point in opposite directions. On OpenRouter — the neutral routing layer where a large share of independent developers actually spend their tokens — the seven highest-volume models all ship open weights, and open models moved from a negligible share of routed tokens in 2024, to roughly a third by late 2025, to the majority by the middle of 2026. Open weights now take 72.4 per cent of the tokens routed to the platform’s top ten models. In Menlo Ventures’ enterprise survey, open models’ share of enterprise LLM usage went the other way: 19 per cent in 2024 down to 11 per cent as of December 2025, with closed models carrying roughly 87 per cent of workloads and the leaderboard of enterprise spend reading Anthropic 32, OpenAI 25, Google 20, Llama 9, DeepSeek 1.
Both numbers are measured honestly. Both describe the same eighteen months. They diverge because they count different things — one counts what developers route, the other counts what enterprises pay for — and citing either without the other produces a confident, useless story. The interesting question is what sits in the gap between them, because that is where every procurement decision in a regulated enterprise is currently being made, mostly without evidence.
A third figure circulates and belongs here rather than in a footnote: AI.cc’s 2026 AI API Infrastructure Report puts open weights at 38 per cent of API tokens in the first quarter of 2026, measured across more than eight thousand API accounts. That is a third denominator again — API traffic, not enterprise workload share and not one platform’s routing — which is why it sits between the two numbers above rather than contradicting either. Three measurements, three populations; the only wrong move is to treat any one of them as the market.
This is the long version of the map: family by family, licence by licence, workload class by workload class, with the places the evidence is thinner than the headline. Several of the numbers below circulate in two incompatible versions, and where that is true I have said which version applies and why — a benchmark index that changed methodology mid-year, a market-share figure whose denominator is not what most people citing it assume, a coding benchmark its most prominent publisher has since retired. That is the whole method. You cannot run a buying discipline on numbers you have not checked, and I would rather show the seams.
The scissors: developer routing and enterprise spend moved in opposite directions
Start with the shape of the divergence, because it is the finding most often reported as a single trend in whichever direction the writer prefers.
On the routing side, the direction is unambiguous. OpenRouter’s June 2026 analysis found that every one of its seven highest-volume models ships open weights, and that open weights had crossed from minority to majority of routed tokens over roughly eighteen months. Secondary reporting puts the platform at something like 25 trillion tokens a week across more than eight million developers, which makes it a large enough sample to be interesting even before you consider that it is one of the few places where the choice is genuinely unconstrained — no enterprise agreement, no procurement cycle, no incumbent vendor relationship, just a price list and a latency budget. Developer survey data points the same way: among developers adding AI to something, 79 per cent report using open models against 71 per cent using closed.
On the spend side, the direction is equally unambiguous and opposite. Menlo’s enterprise figures show open weights losing enterprise share over the same period in which they gained routing share. The dating is load-bearing and worth stating plainly: no 2026 update exists. The source is a survey of 150 technical leaders published on 31 December 2025, which means the most-quoted number in enterprise AI market commentary is seven months old and predates every frontier open release discussed in this piece.
Production conversion is the number that resolves it
The reconciling statistic is not about preference at all. It is about how far a pilot travels. Roughly 53 per cent of open-model projects reach production, against 63 per cent for closed — and the gap widens rather than narrows with company size. Closed-model conversion climbs from about 54 per cent at small companies to about 73 per cent at large ones. Open-model conversion is nearly flat across the same range: 53 to 57 per cent.
That flatness is the whole story. Scale normally helps: bigger companies have more engineers, more platform capability, more money to throw at an integration. When conversion refuses to improve with scale, the obstruction is not capability or budget, it is something structural that gets no cheaper as the organisation gets larger. The reported barriers say what it is: infrastructure cost 27 per cent, security and compliance 26 per cent, maintenance 24 per cent, deployment complexity 23 per cent, and no vendor support 22 per cent. Four of those five are operational obligations that a closed API absorbs on your behalf and a set of weights hands to you.
So the reading is this. Open weights own experimentation and they own token volume. Closed models own enterprise production spend. Neither position is a verdict on capability, and both are true simultaneously — the models are good enough, and most enterprises are not yet organised to run them. The inference providers demonstrate the point from the other side: Fireworks is reported past a billion dollars annualised, five times year on year, and says that more than 95 per cent of the tokens it serves come from models specialised on customer data. Together is reported around $1.15 billion in bookings, Baseten around $600 million. Fine-tuning an open model on proprietary data, served by somebody else’s platform, is the enterprise open-weight use case. Not “we downloaded the weights.” Someone else still runs the GPUs.
The atlas: ten families and what each one is actually for
Kimi K3, released on 27 July 2026 by Moonshot, is the top open model on the aggregate indices as of this writing. It is a 2.8-trillion-parameter mixture of experts with 104 billion parameters active, a one-million-token context, and a 1.56-terabyte checkpoint — a number worth sitting with, because it is the difference between “we can host this” and “we can host this if we buy a rack.” Against GPT-5.6 Sol it posts GPQA 93.5 to 94.1, Terminal-Bench 2.1 at 88.3 to 88.8, and DeepSWE 67.5 to 73.0. On BrowseComp it leads Claude Fable 5, 91.2 against 88.0. Read that spread carefully: on knowledge-heavy reasoning and on terminal-style tool use it is within a point; on hard agentic software engineering it is more than five points behind; on web-browsing agents it is ahead. One model, three different answers, which is the argument for the workload view later on.
GLM-5.2 from Z.ai, released 17 June 2026, is roughly 750 billion parameters with 40 billion active, a one-million-token context, and — this matters — an MIT licence. It sits at 51 on the Artificial Analysis index at $0.447 in and $3.31 out per million tokens. A GLM-5.5 has been discussed for August 2026; it has not shipped, and nothing in this piece assumes it.
DeepSeek V4, 24 April 2026, MIT, is the strongest cost story in the pack. First-party configuration details are thin — the reported shape is a V4-Pro at around 1.6 trillion parameters with 49 billion active and a V4-Flash at roughly 284 billion with 13 billion active, both with a one-million-token context, and I would hold those specifics loosely. The benchmark figures are firmer, with one large caveat I develop later: V4-Pro at 80.6 per cent on SWE-bench Verified is the highest recorded for an open model, with Flash at 79.0 — on a benchmark that OpenAI has since publicly retired as a measure of frontier coding. Flash prices at $0.054 in and $0.242 out per million. OpenRouter’s own framing is the sharpest version of the economics: DeepSeek runs roughly 150 times cheaper than GPT-5.5’s output costs.
Nemotron 3 Ultra from NVIDIA, on Hugging Face since 4 June 2026, is a 550-billion-parameter hybrid with 55 billion active, and it is the most genuinely open release in this entire atlas — not just weights, but training data, reinforcement-learning environments and recipes. If your definition of open has anything to do with reproducibility rather than redistribution, this is the only family that comes close to satisfying it. It also scores 38 on the current Artificial Analysis index, version 4.1, which places it nineteenth among large open-weight models. A figure of around 48 circulates widely and is not wrong so much as obsolete: it comes from the version 4.0 index at the model’s June launch, before version 4.1 shifted weighting towards agentic workloads a fortnight later. The correction matters more than a version number usually would, because the two numbers answer different questions. At 48 the best American open model is in the conversation. At 38 it is not — and the most reproducible open release in the world is also, on current measurement, the least capable family in this atlas.
Gemma 4, Google, 2 April 2026, is the volume play and the quiet licensing story. The family spans E2B and E4B efficiency variants, a 26-billion-parameter mixture-of-experts with 4 billion active, and a dense 31B, covering more than 140 languages. Google moved the licence to Apache 2.0 and abandoned its custom terms — a direction of travel that runs against the market, where most of the movement has been the other way. More than 400 million downloads and over 100,000 community variants make it the default base for small-model fine-tuning.
Mistral 3 is routinely miscited as a 2026 release. It shipped on 2 December 2025. Mistral Large 3 is 675 billion parameters with 41 billion active under Apache 2.0, and by mid-2026 the family sits mid-table on capability. A July 2026 model exists in early access and has not been released. One detail that cuts against the way European AI is usually written about: Mistral’s own announcement carries no explicit sovereignty framing. The sovereignty argument is being made about them, not by them.
Qwen is the most consequential story in the atlas and the one where I am most careful. The download data is solid and remarkable: 942.1 million cumulative downloads as of March 2026 against Llama’s 476.0 million, and Qwen alone at 153.6 million monthly — more than double the combined 71.2 million of the next eight families. Alibaba’s family is, by a wide margin, the most-used open model line in the world. And per secondary trackers, the flagship line has been drifting closed while that was happening: Qwen3.6-27B in April 2026 as the most recent open general-purpose release, Qwen 3.7 Max closed in May, and a 2.4-trillion-parameter Qwen3.8-Max-Preview shown in July with open weights promised but no date and no licence named. I flag the drift narrative as not yet confirmed first-party. If it holds, the most-downloaded open family in the world is being converted into a funnel, and a great many downstream fine-tunes are sitting on a base whose successor may not exist.
Llama is receding. Llama 4, from April 2025, is still the reference release; there is no credible 2026 flagship. Its licence carries a 700-million-monthly-active-user trigger and attribution requirements, and is not an OSI-approved licence. Meta’s 9 per cent of enterprise usage in the December 2025 snapshot is a real position built on a model line that has not advanced in over a year.
gpt-oss deserves a paragraph precisely because there is nothing to report. OpenAI’s gpt-oss-120b and gpt-oss-20b shipped on 5 August 2025 under Apache 2.0 and have had no successor. That is eleven months cold at the time of writing. The absence is itself a finding: the American lab that made the loudest open-weight re-entry has not followed it, in a year when the Chinese labs shipped four frontier-class open families.
MiniMax M3, 1 June 2026, is roughly 428 billion parameters with 23 billion active and a one-million-token context, scoring 59.0 on SWE-Bench Pro, 66.0 on Terminal Bench 2.1 and 74.2 on MCP Atlas, at 44 on the aggregate index and $0.098 in and $1.21 out. It is the clearest example of the tier below the frontier that is nonetheless entirely adequate for most production agent work, at a price that makes the frontier look like a luxury purchase.
One structural observation across all of them: every one of the seven strongest open models is a sparse mixture of experts, and six of the seven are Chinese. The architecture convergence is as complete as the geographic concentration, and both matter for what follows.
The licence is not a footnote — it decides whether your fine-tune is an asset or a liability
“Open-weight” is a marketing category that spans at least three legally distinct things, and the distinction is exactly the sort that gets discovered late, by someone in legal, after the fine-tune is in production.
The permissive tier — genuine Apache 2.0 or MIT, no field-of-use restriction, no revenue trigger — currently holds Gemma 4, Mistral 3, gpt-oss, GLM-5.2 and DeepSeek V4. If you fine-tune one of these, the derivative is yours on ordinary open-licence terms, and your legal review is a morning’s work rather than a negotiation.
The restricted “open weight” tier is where the two most prominent names sit. Kimi K3 ships under a bespoke Kimi K3 Licence that requires a separate agreement above $20 million of model-as-a-service revenue in a rolling twelve months. Llama 4 carries its 700-million-monthly-active-user threshold and attribution obligations. Neither is an OSI-approved open source licence, and Moonshot is notably precise about this: the company describes K3 as open weight and does not call it open source. The distinction is theirs, not a critic’s.
The reproducible tier has one entry. Nemotron 3 Ultra is the only family in this atlas that approaches what an OSI definition would actually require, because it published the training data and the recipes alongside the weights.
For a regulated buyer the practical translation is short. A revenue-triggered licence is a commercial cliff you have to model at the volume you hope to reach, not the volume you have. An attribution-and-threshold licence is a compliance obligation attached to a model artefact, which means someone has to own it after the launch team disbands. And a licence with no published training data means your model-risk documentation has a permanent hole in the provenance section that no amount of downstream evaluation fills. Read the licence before the benchmark. The benchmark changes every quarter; the licence follows you for the life of the deployment.
The gap widened in 2026, which is not what anyone expected
The consensus story of 2024 and 2025 was convergence: open models closing on the closed frontier, a few months behind and gaining. Two independent measurement efforts now say that trend reversed this year.
Epoch AI’s analysis of 29 May 2026 is the strongest single statement in the evidence: since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months — about 8 points on their capability index, with a 90 per cent confidence interval of 7 to 11. The comparable figure for the period from 2023 to October 2025 was roughly three months. The lag did not close. It grew by a third.
Stanford HAI’s AI Index 2026 arrives at the same direction by a different route, putting the closed lead at 3.3 per cent in March 2026 against 0.5 per cent in August 2024. Two framings of that figure circulate, so I would treat the direction as the finding and hold the decimal loosely — but the direction is the same direction.
The most useful characterisation of the year’s shape comes from Maxime Labonne’s late-July review, which describes open models as having nearly caught the closed frontier and then fallen behind as the closed labs released stronger models. That is a specific and testable claim about sequencing rather than trajectory: the open labs did not slow down, the closed labs shipped. On the Artificial Analysis index at version 4.1, scored at maximum reasoning effort, Kimi K3 sits at 57 as the top open model and GLM-5.2 at 51, against a closed frontier led by Claude Opus 5 at 61, with Fable 5 at 60 and GPT-5.6 Sol at 59. Best open to best closed is roughly four points. One caveat belongs with that number rather than beneath it: the reasoning-effort setting is load-bearing, and the same closed model scores 61 at maximum effort and 60 one tier down. A quarter of the gap everyone quotes is a configuration choice, which is worth knowing before you build a slide out of it.
Epoch’s own caveat belongs in the headline rather than the appendix: their estimate may tend to understate the true gap, because open models tend to perform relatively worse on private benchmarks than on public ones. A four-month lag measured on public evaluations is therefore a floor, not a ceiling.
Two reasons to hold every parity number loosely: eval contamination and distillation provenance
If this piece stopped at the benchmark tables it would be doing the thing it criticises. Two problems sit underneath every number above, and neither is resolved.
The first is evaluation contamination, and it is not a suspicion. OpenAI published its reasoning on 23 February 2026, under the title “Why SWE-bench Verified no longer measures frontier coding capabilities,” and the demonstration is worse than leakage in the ordinary sense: frontier models reproduced gold patches and problem details from the task identifier alone, with no repository and no issue text in front of them. That is pre-training contamination, and it means a high score on that benchmark is partly a memory test.
A second finding in the same work is routinely merged with the first, and should not be. Of a frequently-failed subset amounting to 27.6 per cent of the benchmark, 59.4 per cent had flawed test cases that rejected correct patches. Broken tests and contaminated training data are different failures pointing in opposite directions — one inflates scores, the other suppresses them — and the widely circulated shorthand of “59 per cent contaminated” is a merger of two numbers that describes neither. OpenAI’s interim recommendation is the SWE-bench Pro public split, alongside investment in private evaluations.
This lands directly on the most-cited open-model result in this piece: DeepSeek V4-Pro’s 80.6 per cent on SWE-bench Verified. I am not withdrawing the figure — it comes from a routing platform with no stake in the outcome — but it now carries a permanent asterisk, and so does every closed-model score it is compared against. The contamination-resistant alternatives, LiveCodeBench and Terminal-Bench, are what I would run instead, and the shortest honest summary is that the highest recorded open coding score sits on a benchmark its own most prominent user has publicly retired.
The second is distillation provenance, and it is the more interesting problem. Secondary analysis reports that Kimi K3 self-identifies as Claude models, shows a 0.72 per-task correlation with Fable 5, and that its predecessor Kimi-K2 showed 82.7 per cent agentic-behaviour similarity with Sonnet 4.5. None of this is proof of anything by itself — self-identification is a well-known artefact of training on web text saturated with model outputs, and behavioural similarity between models trained on overlapping distributions is expected. But it bears on a question that matters for anyone treating open weights as a sovereignty position: whether the open frontier is independently derived, or whether it is a compressed, redistributable image of the closed frontier. If it is substantially the latter, then the four-month lag is not a race — it is a processing delay, and the open frontier’s ceiling is set by a lab whose weights you cannot see.
I do not know which it is, and neither does anyone publishing on it. What I would say confidently is that the sovereignty argument for open weights is weaker than its advocates think if the capability was distilled from a foreign closed model, and stronger than its critics think if it was not — and that this is an empirical question nobody is currently funded to answer.
Parity is a property of the workload, not a property of the model
Aggregate index scores are the wrong instrument for a buying decision, because they average over workload classes that behave completely differently. Broken out, the picture is much sharper, and much more useful.
Embeddings and retrieval: open has won outright. This is the strongest parity claim in the evidence and the one that gets the least attention. Qwen3-Embedding-8B is reported at 70.6 on MTEB against roughly 64.6 for OpenAI’s embedding model and 68.3 for Google’s, and BGE-M3 is reported as the most-used embedding model in production retrieval systems. These figures come from secondary write-ups rather than first-party cards, but five independent sources agree, which is about as good as secondary evidence gets. The practical consequence is large: the retrieval layer of your stack, which is also the layer that touches the most sensitive data, can run entirely on weights you hold, at no capability cost. If you are going to self-host one thing, host this.
Agentic terminal and tool use: parity at the top. Kimi K3 at 88.3 on Terminal-Bench 2.1 against 88.8 is half a point, which is noise. K3 leads on BrowseComp. For the class of agent work that consists of calling tools, reading their output and deciding what to do next, the open frontier is not behind.
Hard agentic coding: a real, persistent gap. DeepSWE 67.5 against 73.0 is five and a half points and it is not noise. Labonne’s assessment — that closed models still keep a clear edge on real-world agentic coding — matches what the numbers show. Note the tension with the SWE-bench Verified result, where open leads: the benchmark that shows parity is the one with the contamination problem, and the benchmark that shows a gap is the harder, more agentic one. That is not a coincidence to wave away.
Frontier reasoning: closed leads, narrowly. K3’s GPQA sits 0.6 points behind. Broader reasoning gaps of three to eight points appear in secondary analyses that I would not chart without better sourcing.
Multilingual: parity in breadth, unproven in depth. Gemma 4 covering more than 140 languages is a real and verifiable claim about coverage. It is not a claim about per-language quality, and the literature on negative transfer in low-resource languages plus the tokenisation penalty says you should assume depth is worse until you have measured it on your languages. For an Indian or Gulf deployment this is the difference between a marketing number and a deployable one.
Production reliability: closed leads on everything that is not capability. The 53-versus-63 per cent conversion gap is the aggregate expression of tool-schema stability, uptime commitments, scaffolding maturity and someone to call. None of it appears on a leaderboard. All of it appears in a post-incident review.
The price cliff is real, and it is not the reason to self-host
The cost delta between open and closed inference is now large enough to change what is buildable. OpenRouter’s roughly 150-times figure for DeepSeek against GPT-5.5 output pricing is the sharp end; the general shape is DeepSeek V4-Flash at $0.242 per million output tokens, MiniMax M3 at $1.21 and GLM-5.2 at $3.31, against closed frontier models in the $15 to $25 range. Roughly two orders of magnitude of price for something in the region of four index points of capability. Kimi K3, on the most cited framing, is a few points off the top at about a third of the price.
And then the counterintuitive part, which I think is the single most useful thing in this piece for anyone about to buy GPUs.
Secondary analyses — four of them, consistent — put self-hosted inference at $0.17 to $1.00 per million tokens against $5 to $15 for closed APIs, with break-even against a budget open API at around 5.7 billion tokens a month, or against a $5-per-million closed model at around 256 million tokens a month, assuming 60 to 70 per cent GPU utilisation. That assumption is doing enormous work. The same analyses note that a GPU running at 10 per cent utilisation costs roughly ten times as much per token as one running well. Most enterprise inference workloads are bursty, which means most enterprise GPU fleets run nowhere near 70 per cent.
Put those together with a price drop of roughly 50 times in 36 months for GPT-4-class capability — 112 times from GPT-4’s original $45 — and the conclusion inverts the usual pitch. Cheap open-model APIs destroyed the cost case for self-hosting. If your reason for holding weights is that it will be cheaper, run the arithmetic at your actual utilisation, because it very probably will not be.
The reasons that survive the arithmetic are all sovereignty reasons: data that cannot leave a jurisdiction or a network; a model version that cannot change under a production system without your consent; the ability to reconstruct, months later, exactly which weights produced a given action. That last one is not a preference. Under the revised interagency model risk guidance of 17 April 2026 — OCC Bulletin 2026-13 and the parallel Fed and FDIC issuances, which supersede SR 11-7 and SR 21-8 — generative and agentic models are expressly outside scope, with separate guidance promised. Read that as a deferral rather than an exemption: every obligation attached to the underlying action still binds, and what has been removed is the framework that would have specified the controls. Nobody is going to hand you an agent-control specification. The separate guidance, when it comes, will be written against whatever the industry has already built. Holding the weights is the only complete answer anyone has shipped to the question of what happens when a provider updates a model underneath a system you have already validated.
Sovereign programmes are being built on top of open bases, not instead of them
One number carries this entire section. Counterpoint’s Sovereign AI LLM Index, published 30 July 2026, finds that 56 per cent of sovereign large language models are adapted models built on existing open-source bases. National AI capability, as it actually exists in 2026, is overwhelmingly a fine-tuning and continued-pretraining exercise on top of somebody else’s weights — and given the geographic concentration noted earlier, in a great many cases those weights are Chinese.
The same index puts the UAE’s Falcon H1 first globally across all four of its dimensions, with Saudi Arabia’s ALLaM, Jais 2, K2 Think V2, GigaChat 3.1 and India’s Sarvam 105B rated highly. The Gulf’s technical position is not a press release: TII’s Falcon-H1R-7B is a transformer and Mamba2 hybrid with a 262,000-token context, which is a genuine architecture bet rather than a re-labelled fine-tune, and it sits inside a capital environment — G42, AI71, MGX’s $100 billion — that removes compute as the binding constraint.
India’s programme is the clearest case of a state buying the layer beneath the models. The IndiaAI Mission is funded at ₹10,371.92 crore, roughly $1.25 billion, and its distinguishing feature is the compute tender: on the order of 34,000 GPUs available at approximately ₹65 per GPU-hour, with one tracker counting 38,231 — the same order of magnitude either way. That price is the policy. It converts model development from a capital question into an operating one for twelve funded organisations including Sarvam, Soket, Gnani, Gan, Avataar, IIT Bombay’s BharatGen, GenLoop, Zenteiq, Intellihealth, Shodh, Fractal and Tech Mahindra. BharatGen’s Param2, a 17-billion-parameter mixture of experts covering 22 languages, is what that funding produces: not a frontier competitor, and not trying to be — a model that serves languages the global frontier treats as an afterthought.
Europe is the cautionary case rather than the anchor. OpenEuroLLM is funded at €37.4 million across 20 organisations, is a year into a three-year programme, has no flagship model, and describes itself as compute-constrained. Set that against France committing €109 billion to AI infrastructure, Portugal shipping a model for €5.5 million, and Canada’s $890 million programme, and the lesson is not that Europe underfunded AI — it is that consortium structures and infrastructure spending are not substitutes for the thing that actually produces a model, which is concentrated compute pointed at a single team with a deadline. Western Europe and South America are the only regions where closed-model adoption leads, which is what a capability vacuum looks like from the demand side.
The concentration shows up in routing as well, and this is a place where two correct figures are routinely conflated into one wrong story. Chinese models took 72.7 per cent of the open-weight tokens routed through OpenRouter in January 2026 — a share of the open-weight slice alone, and a complete inversion of where China and the United States stood fourteen months earlier. Separately, Chinese models took 46.4 per cent of all routed tokens by mid-2026, against 35.7 per cent for American ones — a share of everything, open and closed together. Those are different denominators, not two points on a trend, and plotting them as a line manufactures a decline that did not happen. No open-only figure exists for mid-2026, which is why the January date has to travel with the 72.7.
The wider frame: more than 70 national AI strategies now exist, 47 countries restrict foreign processing of certain data, and WAICO — convened in Shanghai with 29 member states — includes no major Western democracy. Sovereignty has stopped being a slogan and become a set of procurement constraints that will show up in your requirements documents whether or not your architecture is ready for them.
What is law, what is a proposal, and the tap China is considering closing
This is the area where imprecision is most common and most damaging, so I will state each item at the confidence it deserves.
In the United States, the download bans and licensing regimes under discussion are proposals, not law. No obligation currently attaches to a US enterprise from them. Anyone telling you to plan around a US open-weight export restriction is describing a possible future, and should say so. Separately, and at lower confidence, there are reports that US export controls were applied to a frontier closed model in June 2026 — I flag it because if accurate it means the control regime is landing on closed models too, which changes the shape of the argument, but I would not build on it until it is confirmed first-party.
In China, MOFCOM is consulting on a tiered export regime that could include a ban on the public release of the most capable open models. That is the consequential one. Six of the seven strongest open models are Chinese; the largest supplier of open weights in the world is actively considering closing the tap. Every enterprise architecture that depends on a Chinese open base — which, per Counterpoint’s 56 per cent, includes a majority of the world’s sovereign programmes — has a single-jurisdiction dependency it probably has not written down.
The industry position is on the record and is unusually broad. The letter published 24 July 2026, “Open Weights and American AI Leadership,” carries more than 230 signatories including Amazon, Google, Meta, Microsoft, NVIDIA and OpenAI. It is worth reading for the coalition it represents rather than for its evidence: it makes no quantitative claims, and I have seen figures attributed to it that are not in it.
The most interesting argument in this space comes from Sequoia’s 24 July analysis, and it is uncomfortable in a productive way. Their figures show Qwen’s share of new open fine-tunes rising from 1 per cent in January 2024 to 69 per cent in February 2026, and they argue that a majority of American AI startups now use Chinese open weights somewhere in their stack. The mechanism they identify is the part worth sitting with: the terms of service of the frontier closed labs prohibit training on their outputs, which makes Chinese open weights the only lawful path for capability transfer into a new model. Restricting them does not remove the demand for a base model to build on; it removes the legal one. ATOM’s assessment, published 4 July, is that US frontier open models are within six to twelve months of feasibility. Whether anyone funds them is a different question, and the eleven-month silence from gpt-oss is the current answer.
What this argument does not prove
Four honest limits, because the piece is worth less without them.
The enterprise spend figure is the weakest load-bearing number here. Menlo’s 11 per cent comes from a survey of 150 technical leaders published on 31 December 2025, and no refresh has followed. It is therefore a December 2025 reading being used to characterise mid-2026, in a market where the strongest open releases in this atlas all landed between April and July 2026 — after the fieldwork closed. If a 2026 refresh shows open enterprise share rising, the scissors close and the central framing of this piece needs rewriting. I would rather flag that than defend it later.
Benchmark numbers across labs are not strictly comparable. Harness differences, scaffolding differences and contamination all move scores by more than the gaps being discussed. Every parity or near-parity claim above should be read as “close enough that you must test it yourself,” never as a ranking.
Aggregate capability indices average away the thing you care about. A single index number that puts the open frontier three or four points behind is compatible with open leading on your workload by a wide margin and trailing badly on another. The workload-class section is the useful one; the index section is context.
And the sovereignty case has an unresolved dependency at its foundation. If the distillation-provenance signals hold, then building national capability on open bases inherits a ceiling set by closed labs in another jurisdiction, and the independence being purchased is partly notional. I think the sovereignty argument survives this — control over deployment, data residency and version stability are real regardless of where the capability originated — but it survives in a smaller form than its advocates usually state.
The buyer’s discipline: pin, evaluate, move
What follows from all of this is not a model recommendation. The map redraws quarterly; a recommendation made today expires before your procurement cycle closes. What does not expire is the process, and it is ordinary engineering rather than anything you buy.
Pin a version and write down what you pinned. Not the family, the exact checkpoint. This is the single control that distinguishes a governed deployment from an ungoverned one, and it is the entire practical reason to hold weights rather than call an API. A model that changes under your production system without your consent invalidates every evaluation you ran against it, and the revised model risk guidance’s deferral on agentic systems means no supervisor is going to specify this control for you.
Evaluate on your workload, at your production context shape. Not on aggregate benchmarks, and specifically not on the contaminated ones. Build a parity harness against your pinned baseline, run every candidate release through it, and keep the results. The harness is a week of work and it is the only thing that converts the atlas above into a decision.
Move deliberately, on evidence, when the delta clears a bar you set in advance. Teams that chase every release inherit every regression. Teams that never move pay a compounding capability tax. The discipline that works is the one manufacturing uses for supplier qualification: periodic, boring, evidenced.
Decide self-hosting on sovereignty, not on cost. Run the utilisation arithmetic honestly. If the answer is that hosted open-model APIs are cheaper — and at most enterprise volumes and utilisations they now are — then be clear that you are holding weights for control, residency and version stability, and make the business case on those terms. A self-hosting programme justified on cost that then fails to deliver cost savings gets cancelled, and it takes the governance benefits with it.
And read the licence at the volume you hope to reach. A $20 million revenue trigger is not a constraint on a pilot. It is a constraint on the success case, which is the only case worth planning for.
The aisle is full, the prices have collapsed, and the capability is genuinely there for most of what enterprises are trying to do. What is missing is not models. What is missing is the buying discipline — and the enterprises that build it first will spend the next two years choosing calmly from a market everyone else is chasing.