There is no shortage of open-weight comparison tables. They rank the same models on the same composites and they are, for practical purposes, interchangeable.
None of them carries the column that decides whether you can deploy the model into a decision anyone will later ask you about: *what its publisher actually discloses about what it was trained on.*
That column got more valuable on 1 August, and this piece is mostly about why.
Name the ruler
Before any number: the composite positions below come from one scoring system. *BenchAlign v5.2 tracks 381 benchmarks and weights 27* of them into its composite, across 104 supported and 111 estimated models among 379 tracked.
Change the weighting and you change the leader. The top three are separated by under two points on a composite drawn from 7% of the benchmarks it tracks. That gap is smaller than the difference plausible alternative weightings would produce. Treat the ordering as one defensible view, not as a ranking of the models.
These are reported rankings from an aggregator, not measurements I made. Where a selection decision matters, re-benchmark on your own tasks — the composite tells you where to look, not what to choose.
The matrix
Filter by provenance disclosure level or by deployment footprint. The second facet is the one that changed this month.
Open-weight models, August 2026 — capability alongside provenance
Composite positions are under BenchAlign v5.2 (27 of 381 benchmarks weighted). Provenance records what the publisher discloses, not what is true.
6 of 6 rows
| MiniMax M3 | 68.8 | 2026 | Cluster | Category-level | Publisher describes corpus composition by category. No per-source manifest published. | Leads the BenchLM August composite. Ranking is under BenchAlign v5.2, which weights 27 of 381 tracked benchmarks. | BenchLM open-source leaderboard |
| Hy3 | 67.9 | 2026 | Cluster | Category-level | Corpus described at category level in the model card. Opt-out mechanism not documented. | Second on the same composite, within roughly one point of the leader — a gap smaller than the spread between weighting schemes. | BenchLM open-source leaderboard |
| GLM-5.1 | 66.9 | 2026 | Cluster | Category-level | Corpus categories published. Licensed-corpora list not published. | Third on the August composite. Available through a coding-focused subscription tier as well as open weights. | BenchLM open-source leaderboard |
| Qwen3.8 Max | 60.9 | 2026-08 | Cluster | Category-level | Highest-ranked model released during August 2026. Training corpus described by category. | The strongest August release by composite. Worth re-benchmarking rather than taking the aggregate on trust. | BenchLM open-source leaderboard |
| Kimi K3 | Not on composite | 2026-07-27 | Cluster | Category-level | 2.8T-parameter mixture-of-experts, 1M-token context, native vision. Full open-weight release 27 July 2026. | The long-context option. A 1M-token window changes which workloads need a hosted tier at all. | LLM release tracking |
| Meta 30B agent | Not on composite | 2026-08 | Single GPU | Category-level | 30 billion parameters, runs on a single GPU. Local execution reduces latency and removes a class of data-egress concern. | The architecturally significant release of the month. Dissolves the premise that capable local inference requires a cluster. | August 2026 release reporting |
Every model in this set discloses training data at category level only. That uniformity is the finding — see below.
The uniformity is the finding
Every model here discloses at the same level: *category*. The publisher will tell you the corpus contained code, web text, books, scientific papers. None publishes a per-source manifest. None documents whether an opt-out mechanism was honoured. None lists licensed corpora.
That is not an accusation — category-level disclosure is the current norm and publishing more carries real legal exposure. But it means something practical: *if a buyer asks what your model was trained on, the honest answer for every model above is the same, and it is not a differentiator today.*
Which is exactly why it is about to become one.
The inversion
On *1 August 2026 a new copyright law came into force in Saudi Arabia, with implementing regulations issued the same day. It permits reproduction of original works without the author's permission and without compensation* where the purpose is developing artificial intelligence products and algorithms.
The conventional reading is straightforward: permissive training rules lower build costs, capability migrates toward permissiveness, and provenance becomes a tax paid by whoever happens to be regulated.
That reading holds only if buyers do not care, and buyers increasingly have to — not for ethical reasons but for *attestation* reasons. Deploy a model into a regulated decision and someone will eventually ask what it was trained on. Not out of curiosity: the answer determines whether you can indemnify, whether you can defend an output, whether you can respond to a discovery request, and whether an insurer treats a bad outcome as covered.
So the inversion. When provenance stops being universally required, it stops being a cost of doing business and starts being a claim your competitor cannot make. A scarce, verifiable property that a buyer needs is not a burden — it is a differentiator, and it prices accordingly.
The available strategic error is to read a permissive jurisdiction as permission to stop tracking. The move is the opposite: track deliberately, everywhere, and treat the record as an asset — because the alternative is discovering in eighteen months that your most capable model is the one you can say least about.
The footprint column
The other change worth building against is in the deployment column, and it is easy to miss because it does not move a leaderboard.
*Meta's 30-billion-parameter agent runs on a single GPU.* For two years the sovereign-deployment argument rested on an unstated premise: capable local inference requires a cluster, a cluster requires a programme, and a programme requires a sponsor. That premise did most of the work in every build-versus-rent decision, because it made renting rational for everyone below a certain scale.
A capable agent on one GPU removes it — not universally, and anyone expecting a 30B agent to substitute for a frontier model will be disappointed. But for workloads that are bounded, repetitive, latency-sensitive, or touching data that should not cross a boundary, the arithmetic changed this month.
That matters for different reasons in different places, and this is a global stack rather than a regional story. Where data-residency duties bind, local execution removes the hardest part of the compliance argument rather than mitigating it. Where sovereign-AI programmes fund national capability, a footprint needing no hyperscaler contract changes what capability means operationally. Where cost governance dominates — enterprise agent token cost is now an explicit axis of vendor competition — a fixed-cost local tier beneath a variable-cost API tier is arithmetic. Where latency decides usability, the round trip is the product.
Kimi K3's *one-million-token context window* cuts the other way and is worth holding alongside it: some workloads that looked like they needed a hosted frontier tier needed a long context, which is now available in open weights.
What to record per model
If provenance is an asset, maintain it like one. The fields that a procurement questionnaire actually reaches for:
- *Weights pinned to a digest*, never a tag, and independently hashed rather than trusted.
- *Licence identifier, plus the obligations that survive redistribution* — attribution, notice retention, statement of modifications. This is the field most bills of materials omit and most procurement questionnaires ask about.
- *Disclosure level, stated honestly* — full, partial, or none. An honest "partial, category-level only" is defensible. An implied "clean" is not.
- *Whether an opt-out mechanism was honoured*, including "unknown", which is the correct answer for every model in the table above.
- *An accountable owner and a review date on every gap.* A gap with a name against it is a managed risk; a gap without one is an unexamined one.
- *Evaluation results with their scoring system named*, plus a blind parity result against a pinned baseline.
What this table does not tell you
- *Whether the disclosure is accurate.* The provenance column records what publishers say. Nobody in this set has been independently audited, and this reference does not imply otherwise.
- *Whether the model suits your task.* A composite drawn from 27 weighted benchmarks is a pointer, not a selection. Re-benchmark on your own work.
- *Anything about containment.* Weights carry no controls. The most instructive agent security failure of 2026 happened inside an organisation with total sovereignty over its stack — sovereignty is a supply-chain property and containment is an architectural one.