Open-weight parity evaluation
Whether a model you control can match a hosted one on the specific workflow, measured rather than argued.
The premise
The sovereignty argument is usually made on principle and lost on quality. If an open-weight model inside your boundary is materially worse at the workflow it replaces, sovereignty was bought with a regression nobody measured — and the argument deserves to lose. The only version of this worth making is empirical and per-workflow, because a general benchmark score has never told anyone whether a specific enterprise task will survive the swap.
This strand is running and has produced nothing publishable yet. It is listed so the method can be argued with before there are results to defend.
What this changes for a practitioner
- Whether the sovereign option is viable for your workflow, as a measurement rather than a position
- The residual gap where it is not, stated plainly, so the decision is informed rather than ideological
- A method you can run yourself against your own task, instead of trusting a leaderboard built on someone else's
How it is actually done
- Blind evaluation against a pinned frontier baseline, scored per workflow rather than on a general benchmark
- The task, the rubric and the baseline are fixed before any model is run, so the comparison cannot be tuned after the fact
- Frontier models are used to judge, never to generate the material being judged
- Provenance is recorded for every run — model version, quantisation, prompt and configuration — because an evaluation you cannot reproduce is an anecdote with a number on it
What is unresolved
Published as unresolved on purpose. A practice that only publishes what it has solved is doing marketing.
How much of the measured gap is capability and how much is prompt and scaffolding tuned to one provider's conventions?
Does parity hold under the long-tail cases that matter in regulated work, or only at the median where benchmarks look?
What is the honest total cost of the sovereign option once evaluation, hosting and maintenance are counted — and at what volume does it invert?
Standards, regimes and dated claims on this page were last checked against their primary sources on . Instruments move — several cited here changed inside the last year — so verify against the source before you rely on one. If you find something stale, tell me and I will correct it.
Working on any of this? Compare notes — in the open or under NDA. All strands: research.
Where this meets a live decision
Research is upstream of the advisory work. If this strand maps onto something you are deciding now, that is the useful conversation.