The build-versus-buy meeting has become a fixture of the AI programme, and it is almost always argued at the wrong altitude — whole-solution versus whole-solution, this vendor against that platform against “our engineers could do it.”
At that altitude the debate is unresolvable, because every option is simultaneously right and wrong: the platform is genuinely faster, the in-house build genuinely fits better, and the argument turns on temperament rather than evidence. The productive version of the conversation happens one level down, layer by layer, and it turns on a single test.
Is this layer a commodity that improves when the market competes — or is it the layer where my obligations and my institutional knowledge live?
Run the stack through that test and the answer pattern is surprisingly stable across industries, which is itself the interesting finding. Sourcing debates feel bespoke and mostly are not.
Rent the models — ruthlessly
Model capability is a supplier market with every property you want in one: evenly distributed across several credible providers, improving quarterly, and collapsing in price. Blended enterprise token cost fell 67 percent year on year — from $18.40 to $6.07 per million, across an analysis of 2.4 billion API calls — driven by open-weight competition and multi-model routing.
Renting is straightforwardly correct here, and the reasoning is not primarily about cost. It is that a layer improving this fast is a layer where any commitment you make is a bet against the market, and the market has been winning that bet every quarter for three years.
The condition attached is the part programmes skip: renting is correct provided you preserve the ability to switch. Which means your prompts, your evaluations and your fine-tunes must not be denominated in one vendor's currency. Prompts written against one model's quirks, an evaluation harness that only runs through one provider's API, fine-tunes that exist only as a vendor-side artefact — each is a small, sensible-at-the-time decision that converts a rental into a dependency.
Rent the model. Never rent your portability. And the test for whether you have kept it is not whether you could switch in principle — it is whether anyone has actually run your evaluation suite against a second provider in the last quarter. Portability that has never been exercised is a belief, not a capability.
Buy the plumbing
Orchestration frameworks, vector infrastructure, monitoring, gateways: competitive markets, real products, no institutional knowledge inside them. Buying is faster than building, and switching costs are manageable if you keep your data formats portable.
This is the layer where engineering teams most often argue for building, and where the argument is most often wrong. The reasoning usually runs that the available products do not quite fit, which is true and is not decisive — the market improves this layer continuously and will keep improving it after your team has moved on to something else. An in-house version fits perfectly on the day it ships and degrades relative to the market every quarter thereafter.
The one diligence question that matters here is the same one the agent-washing filter turns on: does the component respect an authority model, or fight it? A gateway that assumes it holds the entitlement logic, or an orchestration framework with no concept of an action being refused, is not a neutral piece of plumbing. It is a component that will push back against the two layers you are about to build, and you will discover that in month seven.
Build the two layers that are you
Two layers fail the commodity test on both grounds simultaneously, and they are the ones worth owning.
The context layer. Your knowledge, under your entitlements, with your provenance. No vendor can supply this ready-made because it is made of your organisation's specifics — which systems hold what, who may see it, what the retention obligations are, how the entitlement structure actually works rather than how the org chart says it does. A vendor can sell you the retrieval machinery. They cannot sell you the answer to what this agent may know.
The evidence layer. The attested record of what your systems knew and did. This is where regulatory obligation attaches directly: the DPDP Act's purpose limits, the obligations left standing after OCC Bulletin 2026-13 withdrew the framework that would have specified the controls, every supervisor's version of prove it. A purchased logging product can hold the records. It cannot be the thing that owes them.
These two are also, not coincidentally, where switching rights live. Own them and every rented layer above stays rented on your terms — because the thing that makes a vendor hard to leave is rarely the capability, which is replaceable, and almost always the accumulated context and history that only exist inside their system.
The red flag that overrides everything
In your next review, notice who answers the governance questions.
If what your agent may do, what it can see, and what record it leaves are all answered by the vendor, the sourcing decision has already been made — by default, in the wrong direction, and usually without anyone noticing they were making it.
Capability can be rented. Accountability cannot. Your regulator holds you, not your platform; a contract clause transfers blame at most, never obligation. This is not a legal technicality, it is the structural fact the whole framework rests on: there is no commercial arrangement that makes someone else the answerable party for what your systems did to your customers.
Which is why the red flag is diagnostic rather than merely annoying. A programme where the vendor answers the governance questions is not a programme with a vendor-management problem. It is a programme that has quietly relocated its accountability to somewhere it cannot legally go.
Three sourcing errors to name at the table
Each of these is individually defensible in the meeting where it is made, which is exactly why they need naming in advance.
Building commodity. An in-house orchestration framework is an expensive way to make future hires sad. Engineering pride is not a moat. The cost is not only the build — it is the attention that was owed to the two layers nobody else can build for you, spent instead on a layer three vendors are competing to give you.
Renting identity. Letting a platform hold your entitlement model “for convenience” converts your authority structure into their configuration data. You will meet it again, priced as switching cost. This is the quietly most expensive of the three, because it is invisible until you try to leave — at which point the question is not whether you can afford the new platform but whether you can afford to re-derive who is allowed to do what.
Buying governance. Compliance-in-a-box products can implement your controls but cannot be your controls. A purchased policy engine enforcing policies nobody in the building decided is theatre with a licence fee. The tool is frequently good; the substitution is the error — buying the enforcement mechanism is sensible, buying it instead of deciding the policy is not.
The sequencing implication
Here is the practical payoff, and it is the part I would put on the board slide.
Build the two thin layers first. An entitlement model and an evidence pipeline are weeks of work, not years — thin by construction, because they encode decisions rather than implementing capability. And once they exist, every subsequent rent-or-buy decision becomes low-stakes, because nothing above those layers can hold you hostage.
Programmes that sequence the other way arrive at year two with impressive rented capability, no owned accountability, and a renewal negotiation they cannot afford to lose. Every decision along that path was defensible on the day it was made. The failure is not in any of the decisions; it is in their order.
That is why this is not a budget question. The two layers are cheap. What makes them hard is that they produce nothing demonstrable in the first month, while renting capability produces a demo — and the programme's early credibility usually depends on the demo. Naming that trade-off explicitly at the start is most of the work.
The layer the test keeps missing: evaluations
One layer sits awkwardly in the framework and deserves calling out, because programmes consistently misfile it.
Evaluations look like plumbing. There are products, the market is competitive, and building a harness in-house feels like the commodity error. So teams buy an evaluation platform, and the decision seems obviously right.
But run the actual test on it. Is an evaluation suite a commodity that improves when the market competes, or is it where institutional knowledge lives? The machinery is commodity — runners, dashboards, statistical treatment, regression tracking. The suite itself is the opposite: it is the accumulated encoding of what your organisation has decided good looks like, which failure modes you have been burned by, and which edge cases matter in your domain. That is institutional knowledge in its purest form, and it compounds.
So evaluations split across the line rather than sitting on one side of it, and the sourcing answer is correspondingly split: buy the harness, own the suite. Concretely, that means your test cases, your grading criteria and your accumulated failure library live in your repository in a portable format, and the platform executes them. A programme whose evaluation suite exists only inside a vendor's product has rented the one artefact that would have told it whether switching was safe — which is a particularly unfortunate thing to have rented, because it is the instrument you would need to exercise the portability you thought you had preserved.
This is the general shape of the awkward cases, and it is worth stating as a rule: where a layer has both machinery and content, the machinery is usually commodity and the content usually is not. Apply the test to the content, not to the category.
How to tell whether you have already rented your identity
Renting identity is the error that has usually already happened by the time anyone asks the question, so it needs a diagnostic rather than a warning.
Four questions, answerable in an afternoon by someone with access.
If you terminated your largest AI platform contract tomorrow, where would the answer to who may do what live? If the honest answer is “in their console”, the entitlement model is theirs. Not contractually — practically, which is the sense that binds when you are trying to leave.
Can you produce, today, a list of what each agent is permitted to reach, in business terms, without logging into a vendor system? A programme that owns its authority model can. A programme that has rented it will produce a screenshot.
When someone's role changes, what has to happen for the agents acting on their behalf to change scope — and who performs that action? If the answer routes through a vendor's configuration UI rather than your identity infrastructure, your authority structure has become their configuration data, exactly as described above.
And the one that usually settles it: if a regulator asked for the authority basis of a specific action from eight months ago, would the reconstruction require vendor cooperation? Needing a supplier's goodwill to answer a supervisor is a position worth discovering before you are in it.
None of these is a reason to leave a platform. They are a reason to know what you would be re-deriving, and to start re-deriving the cheap parts now rather than the expensive parts later.
Running the review as a session, not a project
The layer-by-layer sourcing review is a whiteboard session, not a quarter, and treating it as a project is its own failure mode — a six-week vendor-comparison exercise produces a matrix rather than a decision.
The session works like this. List the layers of the stack you actually have, not the reference architecture — usually six to nine boxes. For each, ask the one test question and force a single-word answer: rent, buy, or build. Where the room disagrees, the disagreement is almost always about whether institutional knowledge lives in that layer, and that is a factual question someone can go and check rather than a matter of preference.
Then do the pass that produces the value: for every layer marked rent or buy, ask what would have to be true for us to switch this in ninety days. The answers are the portability requirements, and they are far more actionable than a general commitment to avoid lock-in. Most of them turn out to be small — an export format, an evaluation suite held locally, a permission model that does not live in someone's console.
Finish by checking the order. If anything in the build column is scheduled after something in the rent or buy column, ask why, and be suspicious of the answer. That single question has changed more programme plans in my experience than any amount of analysis of the vendors themselves.
Three markets, three different pressures
The framework holds everywhere; which error dominates does not.
In North America, the common failure is buying governance. The compliance-product market is mature, budgets exist, and purchasing a policy engine is an easy thing to show a board. The error is subtle because the tools are genuinely good — what gets skipped is the deciding, and a control nobody in the building decided is not a control.
In India, speed and price pressure push hard toward renting everything, including identity. Platforms that offer to hold the entitlement model are offering to remove work from a team that is under-resourced and moving fast, which makes the offer very hard to refuse. This is the market where the ninety-day-switch question earns the most, because the cost of the default is highest and the least visible.
In the Gulf, new-build and integrated procurement dominates, which creates the opposite risk: the whole stack arrives as one decision, so the layer test never gets applied at all. The discipline that matters here is insisting on the layer-by-layer pass even when the commercial arrangement is a single contract — because the sourcing question is about where the obligations live, and that does not change because the invoice is consolidated.
The strongest objection: this is a governance consultant's stack, not a builder's
The serious counter-argument comes from people shipping fast, and it deserves stating properly.
It runs: this framework optimises for a hypothetical future audit rather than for getting something valuable into production. Most AI programmes die of irrelevance, not of governance failure. Building an entitlement model and an evidence pipeline before you know whether the agent is useful is exactly the premature-infrastructure mistake that kills projects — you will build the wrong abstractions, because you do not yet know what the system does. Rent everything, prove value, and retrofit governance once you know what you are governing.
The premise is right and the conclusion does not follow. Most programmes do die of irrelevance, and building elaborate infrastructure before proving value is a genuine failure mode. But the retrofit assumption is where it breaks, for a specific reason: the entitlement model is not infrastructure, it is a set of decisions about who may do what. Those decisions determine what the agent can safely be built to do, so deferring them does not postpone the work — it means the pilot is built against an implicit authority model that nobody chose, and which usually turns out to be “whatever the service account could reach.”
Retrofitting then means rebuilding, because the capability was shaped around permissions it should never have had. That is the actual cost of the deferral, and it is paid at exactly the moment the programme has succeeded and wants to scale — which is the worst possible moment to discover it.
The honest synthesis: the objection is correct that these layers should be thin at the pilot stage, and it is wrong that they should be absent. A one-page entitlement model and a log of what-was-known-at-action-time cost days, not quarters. What should not happen at pilot stage is the full platform build — and nothing in this argument asks for one.
Where this argument is weakest
Two admissions.
First, the boundary between plumbing and the context layer is genuinely blurry, and I have drawn it more cleanly than reality supports. Permission-aware retrieval sits on the line: the retrieval machinery is commodity, the permission model is not, and they are not always separable in available products. Programmes will have to make a judgement there, and the honest guidance is to keep the permission decisions in artefacts you own even when the machinery evaluating them is bought.
Second, the 67 percent price decline is a blended figure reflecting both genuine price cuts and a mix shift toward cheaper models. It is the right number for the claim being made — that this layer is commoditising and multi-year commitments are unwise — but it is not the same as saying any specific model got 67 percent cheaper, and it should not be quoted that way.
The whole thing in three lines
Rent what improves when markets compete. Buy what is genuinely product. Build the thin pair that makes you you — and makes everything else replaceable.
Then check the order. Because the programmes that get this wrong rarely get any single decision wrong; they get the sequence wrong, and find out two years later, in a negotiation.