The build-versus-buy meeting has become a fixture of the AI programme, and it is almost always argued at the wrong altitude — whole-solution versus whole-solution, this vendor against that platform against “our engineers could do it.”

At that altitude the debate is unresolvable, because every option is simultaneously right and wrong: the platform is genuinely faster, the in-house build genuinely fits better, and the argument turns on temperament rather than evidence. The productive version of the conversation happens one level down, layer by layer, and it turns on a single test.

Is this layer a commodity that improves when the market competes — or is it the layer where my obligations and my institutional knowledge live?

Run the stack through that test and the answer pattern is surprisingly stable across industries, which is itself the interesting finding. Sourcing debates feel bespoke and mostly are not.

Run every layer through one test Is this a commodity that improves when the market competes — or is it the layer where my obligations and my institutional knowledge live? RENT ruthlessly Models A supplier market: evenly distributed, quarterly-improving, price-collapsing. Blended enterprise token cost −67% YoY: $18.40 → $6.07 per million (2.4bn calls). Condition: prompts, evals and fine-tunes not denominated in one vendor's currency. BUY it's product The plumbing Orchestration · vector infrastructure · monitoring · gateways. Competitive markets, real products, no institutional knowledge inside them. Condition: keep data formats portable. Ask if it respects an authority model, or fights it. BUILD this is you The context layer Your knowledge, under your entitlements, with your provenance. No vendor can supply it ready-made — it is made of your organisation's specifics. BUILD this is you The evidence layer The attested record of what your systems knew and did. Where regulatory obligation attaches — and where switching rights live. Rent what improves when markets compete. Buy what is genuinely product. Build the thin pair that makes you you — and makes everything else replaceable.

Rent the models — ruthlessly

Model capability is a supplier market with every property you want in one: evenly distributed across several credible providers, improving quarterly, and collapsing in price. Blended enterprise token cost fell 67 percent year on year — from $18.40 to $6.07 per million, across an analysis of 2.4 billion API calls — driven by open-weight competition and multi-model routing.

Renting is straightforwardly correct here, and the reasoning is not primarily about cost. It is that a layer improving this fast is a layer where any commitment you make is a bet against the market, and the market has been winning that bet every quarter for three years.

The condition attached is the part programmes skip: renting is correct provided you preserve the ability to switch. Which means your prompts, your evaluations and your fine-tunes must not be denominated in one vendor's currency. Prompts written against one model's quirks, an evaluation harness that only runs through one provider's API, fine-tunes that exist only as a vendor-side artefact — each is a small, sensible-at-the-time decision that converts a rental into a dependency.

Rent the model. Never rent your portability. And the test for whether you have kept it is not whether you could switch in principle — it is whether anyone has actually run your evaluation suite against a second provider in the last quarter. Portability that has never been exercised is a belief, not a capability.

Buy the plumbing

Orchestration frameworks, vector infrastructure, monitoring, gateways: competitive markets, real products, no institutional knowledge inside them. Buying is faster than building, and switching costs are manageable if you keep your data formats portable.

This is the layer where engineering teams most often argue for building, and where the argument is most often wrong. The reasoning usually runs that the available products do not quite fit, which is true and is not decisive — the market improves this layer continuously and will keep improving it after your team has moved on to something else. An in-house version fits perfectly on the day it ships and degrades relative to the market every quarter thereafter.

The one diligence question that matters here is the same one the agent-washing filter turns on: does the component respect an authority model, or fight it? A gateway that assumes it holds the entitlement logic, or an orchestration framework with no concept of an action being refused, is not a neutral piece of plumbing. It is a component that will push back against the two layers you are about to build, and you will discover that in month seven.

Build the two layers that are you

Two layers fail the commodity test on both grounds simultaneously, and they are the ones worth owning.

The context layer. Your knowledge, under your entitlements, with your provenance. No vendor can supply this ready-made because it is made of your organisation's specifics — which systems hold what, who may see it, what the retention obligations are, how the entitlement structure actually works rather than how the org chart says it does. A vendor can sell you the retrieval machinery. They cannot sell you the answer to what this agent may know.

The evidence layer. The attested record of what your systems knew and did. This is where regulatory obligation attaches directly: the DPDP Act's purpose limits, the obligations left standing after OCC Bulletin 2026-13 withdrew the framework that would have specified the controls, every supervisor's version of prove it. A purchased logging product can hold the records. It cannot be the thing that owes them.

These two are also, not coincidentally, where switching rights live. Own them and every rented layer above stays rented on your terms — because the thing that makes a vendor hard to leave is rarely the capability, which is replaceable, and almost always the accumulated context and history that only exist inside their system.

The red flag that overrides everything

In your next review, notice who answers the governance questions.

Capability can be rented. Accountability cannot. Your regulator holds you, not your platform. What a vendor can supply Model capability Orchestration Monitoring Gateways Infrastructure All rentable or purchasable. All improve when the market competes. None of it carries your obligations. a contract clause transfers blame at most — never obligation What stays with you regardless “Who permitted this action?” Reconstructing what the system knew when it acted Purpose limitation under data-protection law The obligations left standing after OCC Bulletin 2026-13 withdrew the framework that would have specified the controls None of this is procurable. The red flag that overrides everything In your next review, notice who answers the governance questions. If what your agent may do, what it can see, and what record it leaves are all answered by the vendor — the sourcing decision has already been made. By default, in the wrong direction.

If what your agent may do, what it can see, and what record it leaves are all answered by the vendor, the sourcing decision has already been made — by default, in the wrong direction, and usually without anyone noticing they were making it.

Capability can be rented. Accountability cannot. Your regulator holds you, not your platform; a contract clause transfers blame at most, never obligation. This is not a legal technicality, it is the structural fact the whole framework rests on: there is no commercial arrangement that makes someone else the answerable party for what your systems did to your customers.

Which is why the red flag is diagnostic rather than merely annoying. A programme where the vendor answers the governance questions is not a programme with a vendor-management problem. It is a programme that has quietly relocated its accountability to somewhere it cannot legally go.

Three sourcing errors to name at the table

Each of these is individually defensible in the meeting where it is made, which is exactly why they need naming in advance.

Three sourcing errors to name at the table Each is individually reasonable in the meeting where it is made. Building commodity An in-house orchestration framework. An expensive way to make future hires sad. Engineering pride is not a moat. The market improves this layer faster than you can, and keeps doing it after your team moves on. Costs: money, and the attention you owed elsewhere. Renting identity Letting a platform hold your entitlement model — “for convenience”. Converts your authority structure into their configuration data. You will meet it again, priced as switching cost. The most quietly expensive of the three, because it is invisible until you try to leave. Buying governance Compliance-in-a-box. Such products can implement your controls. They cannot be your controls. A purchased policy engine enforcing policies nobody in the building decided is theatre with a licence fee. The tool is often good. The substitution is the error. None of these looks like a mistake at the time. Each is defensible in isolation, which is exactly why they have to be named before the meeting, not after it.

Building commodity. An in-house orchestration framework is an expensive way to make future hires sad. Engineering pride is not a moat. The cost is not only the build — it is the attention that was owed to the two layers nobody else can build for you, spent instead on a layer three vendors are competing to give you.

Renting identity. Letting a platform hold your entitlement model “for convenience” converts your authority structure into their configuration data. You will meet it again, priced as switching cost. This is the quietly most expensive of the three, because it is invisible until you try to leave — at which point the question is not whether you can afford the new platform but whether you can afford to re-derive who is allowed to do what.

Buying governance. Compliance-in-a-box products can implement your controls but cannot be your controls. A purchased policy engine enforcing policies nobody in the building decided is theatre with a licence fee. The tool is frequently good; the substitution is the error — buying the enforcement mechanism is sensible, buying it instead of deciding the policy is not.

The sequencing implication

Here is the practical payoff, and it is the part I would put on the board slide.

Build the two thin layers first. An entitlement model and an evidence pipeline are weeks of work, not years — thin by construction, because they encode decisions rather than implementing capability. And once they exist, every subsequent rent-or-buy decision becomes low-stakes, because nothing above those layers can hold you hostage.

Sequencing is the whole payoff Same decisions. Opposite order. Entirely different position at year two. Build the thin pair first entitlement model + evidence pipeline weeks, not years rent models low stakes buy plumbing low stakes swap a vendor low stakes renegotiate from strength Nothing above those two layers can hold you hostage — because the switching rights were owned before anything was rented. The other way round rent capability it works! ship it buy more platform momentum governance “later” phase two year two impressive rented capability · no owned accountability Every decision was defensible on the day it was made. And the renewal negotiation is one you cannot afford to lose. The pair is thin. An entitlement model and an evidence pipeline are weeks of work. So this is not a question of budget. It is a question of order. Build the thin pair first, and every subsequent sourcing decision becomes low-stakes.

Programmes that sequence the other way arrive at year two with impressive rented capability, no owned accountability, and a renewal negotiation they cannot afford to lose. Every decision along that path was defensible on the day it was made. The failure is not in any of the decisions; it is in their order.

That is why this is not a budget question. The two layers are cheap. What makes them hard is that they produce nothing demonstrable in the first month, while renting capability produces a demo — and the programme's early credibility usually depends on the demo. Naming that trade-off explicitly at the start is most of the work.

The layer the test keeps missing: evaluations

One layer sits awkwardly in the framework and deserves calling out, because programmes consistently misfile it.

Evaluations look like plumbing. There are products, the market is competitive, and building a harness in-house feels like the commodity error. So teams buy an evaluation platform, and the decision seems obviously right.

But run the actual test on it. Is an evaluation suite a commodity that improves when the market competes, or is it where institutional knowledge lives? The machinery is commodity — runners, dashboards, statistical treatment, regression tracking. The suite itself is the opposite: it is the accumulated encoding of what your organisation has decided good looks like, which failure modes you have been burned by, and which edge cases matter in your domain. That is institutional knowledge in its purest form, and it compounds.

So evaluations split across the line rather than sitting on one side of it, and the sourcing answer is correspondingly split: buy the harness, own the suite. Concretely, that means your test cases, your grading criteria and your accumulated failure library live in your repository in a portable format, and the platform executes them. A programme whose evaluation suite exists only inside a vendor's product has rented the one artefact that would have told it whether switching was safe — which is a particularly unfortunate thing to have rented, because it is the instrument you would need to exercise the portability you thought you had preserved.

This is the general shape of the awkward cases, and it is worth stating as a rule: where a layer has both machinery and content, the machinery is usually commodity and the content usually is not. Apply the test to the content, not to the category.

How to tell whether you have already rented your identity

Renting identity is the error that has usually already happened by the time anyone asks the question, so it needs a diagnostic rather than a warning.

Four questions, answerable in an afternoon by someone with access.

If you terminated your largest AI platform contract tomorrow, where would the answer to who may do what live? If the honest answer is “in their console”, the entitlement model is theirs. Not contractually — practically, which is the sense that binds when you are trying to leave.

Can you produce, today, a list of what each agent is permitted to reach, in business terms, without logging into a vendor system? A programme that owns its authority model can. A programme that has rented it will produce a screenshot.

When someone's role changes, what has to happen for the agents acting on their behalf to change scope — and who performs that action? If the answer routes through a vendor's configuration UI rather than your identity infrastructure, your authority structure has become their configuration data, exactly as described above.

And the one that usually settles it: if a regulator asked for the authority basis of a specific action from eight months ago, would the reconstruction require vendor cooperation? Needing a supplier's goodwill to answer a supervisor is a position worth discovering before you are in it.

None of these is a reason to leave a platform. They are a reason to know what you would be re-deriving, and to start re-deriving the cheap parts now rather than the expensive parts later.

Running the review as a session, not a project

The layer-by-layer sourcing review is a whiteboard session, not a quarter, and treating it as a project is its own failure mode — a six-week vendor-comparison exercise produces a matrix rather than a decision.

The session works like this. List the layers of the stack you actually have, not the reference architecture — usually six to nine boxes. For each, ask the one test question and force a single-word answer: rent, buy, or build. Where the room disagrees, the disagreement is almost always about whether institutional knowledge lives in that layer, and that is a factual question someone can go and check rather than a matter of preference.

Then do the pass that produces the value: for every layer marked rent or buy, ask what would have to be true for us to switch this in ninety days. The answers are the portability requirements, and they are far more actionable than a general commitment to avoid lock-in. Most of them turn out to be small — an export format, an evaluation suite held locally, a permission model that does not live in someone's console.

Finish by checking the order. If anything in the build column is scheduled after something in the rent or buy column, ask why, and be suspicious of the answer. That single question has changed more programme plans in my experience than any amount of analysis of the vendors themselves.

Three markets, three different pressures

The framework holds everywhere; which error dominates does not.

In North America, the common failure is buying governance. The compliance-product market is mature, budgets exist, and purchasing a policy engine is an easy thing to show a board. The error is subtle because the tools are genuinely good — what gets skipped is the deciding, and a control nobody in the building decided is not a control.

In India, speed and price pressure push hard toward renting everything, including identity. Platforms that offer to hold the entitlement model are offering to remove work from a team that is under-resourced and moving fast, which makes the offer very hard to refuse. This is the market where the ninety-day-switch question earns the most, because the cost of the default is highest and the least visible.

In the Gulf, new-build and integrated procurement dominates, which creates the opposite risk: the whole stack arrives as one decision, so the layer test never gets applied at all. The discipline that matters here is insisting on the layer-by-layer pass even when the commercial arrangement is a single contract — because the sourcing question is about where the obligations live, and that does not change because the invoice is consolidated.

The strongest objection: this is a governance consultant's stack, not a builder's

The serious counter-argument comes from people shipping fast, and it deserves stating properly.

It runs: this framework optimises for a hypothetical future audit rather than for getting something valuable into production. Most AI programmes die of irrelevance, not of governance failure. Building an entitlement model and an evidence pipeline before you know whether the agent is useful is exactly the premature-infrastructure mistake that kills projects — you will build the wrong abstractions, because you do not yet know what the system does. Rent everything, prove value, and retrofit governance once you know what you are governing.

The premise is right and the conclusion does not follow. Most programmes do die of irrelevance, and building elaborate infrastructure before proving value is a genuine failure mode. But the retrofit assumption is where it breaks, for a specific reason: the entitlement model is not infrastructure, it is a set of decisions about who may do what. Those decisions determine what the agent can safely be built to do, so deferring them does not postpone the work — it means the pilot is built against an implicit authority model that nobody chose, and which usually turns out to be “whatever the service account could reach.”

Retrofitting then means rebuilding, because the capability was shaped around permissions it should never have had. That is the actual cost of the deferral, and it is paid at exactly the moment the programme has succeeded and wants to scale — which is the worst possible moment to discover it.

The honest synthesis: the objection is correct that these layers should be thin at the pilot stage, and it is wrong that they should be absent. A one-page entitlement model and a log of what-was-known-at-action-time cost days, not quarters. What should not happen at pilot stage is the full platform build — and nothing in this argument asks for one.

Where this argument is weakest

Two admissions.

First, the boundary between plumbing and the context layer is genuinely blurry, and I have drawn it more cleanly than reality supports. Permission-aware retrieval sits on the line: the retrieval machinery is commodity, the permission model is not, and they are not always separable in available products. Programmes will have to make a judgement there, and the honest guidance is to keep the permission decisions in artefacts you own even when the machinery evaluating them is bought.

Second, the 67 percent price decline is a blended figure reflecting both genuine price cuts and a mix shift toward cheaper models. It is the right number for the claim being made — that this layer is commoditising and multi-year commitments are unwise — but it is not the same as saying any specific model got 67 percent cheaper, and it should not be quoted that way.

The whole thing in three lines

Rent what improves when markets compete. Buy what is genuinely product. Build the thin pair that makes you you — and makes everything else replaceable.

Then check the order. Because the programmes that get this wrong rarely get any single decision wrong; they get the sequence wrong, and find out two years later, in a negotiation.