The three deployment models
Almost every conversation about data residency collapses into one of these three, and a great deal of confusion comes from suppliers using the word private for all of them.
| Model | Where the data goes | Typical cost shape | Satisfies |
|---|---|---|---|
| Hosted model interface | To the provider, processed in their region of choice unless restricted | Per item. No capital outlay | Most commercial confidentiality requirements, with a data processing agreement |
| Dedicated cloud region or private network | Stays in a named jurisdiction, on infrastructure the provider operates | Per item, plus a platform or commitment fee | Residency clauses that specify a country but not ownership of the hardware |
| On your own hardware | Never leaves your network | Capital cost up front, then power and maintenance | Clauses that require the data to remain on infrastructure you control, and air-gapped environments |
What running it yourself actually involves
An open-weight model is a file. It is downloaded, placed on a server with a suitable graphics processor, and served through a local interface that the rest of your systems call in the same way they would call a hosted provider. From the point of view of the applications sitting on top, almost nothing changes. What changes is that you now own the uptime, the patching, the capacity and the upgrade path.
When residency genuinely forces it
This is worth separating from preference, because on-premise costs more and should be chosen for a reason that can be written down.
A contract says so
Client agreements, particularly in defence, government supply chains and parts of financial services, frequently require that data is not transmitted to a third party processor. This is the most common genuine trigger, and it is a document you can point at.
The data cannot lawfully leave the country
Several jurisdictions, including a number across the Gulf, place residency requirements on personal, health or government-related records. A dedicated region sometimes satisfies this, and sometimes the requirement is explicitly about who operates the infrastructure.
The material is confidential in a way that survives a contract
Unreleased financial results, live legal matters, merger and acquisition work and clinical records are cases where organisations reasonably decline to rely on a processing agreement, however well drafted.
There is no outbound connectivity
Industrial sites, secure facilities and some field operations run without a route to the public internet at all. This is the one case where the decision requires no debate.
The cost, honestly
On-premise trades a per item charge for a capital cost and an ongoing obligation. The crossover depends almost entirely on volume.
| Hosted model | On your own hardware | |
|---|---|---|
| Up front | Nothing | USD 6,000 to 15,000 for a single professional GPU server, and USD 30,000 and beyond for higher concurrency |
| Per item | Fractions of a cent to a few cents, depending on model and length | Effectively nothing once the hardware is bought |
| Ongoing | Usage only | Power, rack space or floor space, patching, monitoring and eventual replacement |
| Who fixes it at 3am | The provider | You, or whoever holds your support agreement |
| Model upgrades | Automatic, sometimes whether you wanted them or not | A deliberate project each time, which is an advantage for stability and a cost in effort |
| Best suited to | Low and medium volume, and anything needing the strongest reasoning | High steady volume, and anything where the data cannot leave |
The commercial argument that is often missed: an on-premise build is a fixed cost that does not grow with the size of your team. Per seat licensing for commercial assistants scales with headcount, so at forty or fifty users the comparison starts to look very different from the way it looks at five. That is one of the few cases where privacy and cost point the same way.
The capability trade-off
Open-weight models are not equivalent to the strongest hosted models and it is unhelpful to pretend otherwise. The gap is uneven, which is the useful part.
- Extraction, classification, routing, summarising and routine drafting run close to parity. These are also the great majority of business automation workloads.
- Long multi-step reasoning, difficult code and unusual analytical work are where the leading hosted models remain measurably ahead.
- Speed on your own hardware is a function of what you bought. A single professional GPU serving a small team is comfortable, and the same hardware serving two hundred users is not.
- The gap has narrowed every year and there is no reason to assume it stops. A workload that is marginal today is a reasonable candidate for review in twelve months.
- The only test that matters is your own. Run fifty of your real, hardest cases through both and compare the output, rather than relying on any published claim, ours included.
The middle path most organisations end up at
Few builds are entirely one or the other. Two patterns cover most of what we see in practice.
Route by sensitivity
Confidential records are processed locally and everything else goes to a hosted model. The routing rule is written down, testable and auditable, which matters more than where any individual item ends up.
Redact, then send
A local model strips names, account numbers and other identifying data, and only the redacted text leaves the network. This gives most of the capability of a frontier model with a much smaller disclosure. It requires the redaction step itself to be tested properly, because a leak here is silent.
Six questions before committing to on-premise
- Which specific clause, regulation or constraint requires it, in writing.
- Whether a dedicated region would satisfy that clause, since it usually costs far less.
- How many people and how many items a day the system has to serve, because that determines the hardware and therefore most of the cost.
- Who patches, monitors and upgrades it, named before the hardware is ordered.
- What the fallback is when the server is down, because a local deployment has no provider status page and no automatic failover unless somebody built one.
- How the model will be evaluated at upgrade time, so that a newer version cannot quietly regress on the cases you depend on.
Our own position, marked as ours
We build both, and we do not treat on-premise as the default. Where a contract, a regulation or an absent network connection forces it, the models run on hardware you own, nothing leaves the premises, and the build is a fixed cost rather than a licence that grows with your team. Where it does not, we say so, because recommending hardware nobody needed is an expensive way to be cautious. That figure of one fixed build instead of a per seat licence is typical of a first build rather than a measured result for a named client.
Frequently asked questions
Is a private cloud region the same as on-premise?
No, although it is often good enough. A dedicated region or a virtual private cloud keeps data inside a jurisdiction and outside the model provider's training pipeline, but the infrastructure still belongs to somebody else. On-premise means the hardware is in your building and the network boundary is yours. Only the second satisfies the strictest contractual clauses.
How far behind are open-weight models?
On extraction, classification, summarising and routine drafting they are close to parity and the difference rarely shows in production output. On long multi-step reasoning and difficult code the leading hosted models remain measurably ahead. The practical approach is to test your own hardest cases against both rather than trust either claim.
What hardware is actually needed?
For a single team's workload, one server with a professional GPU is usually sufficient, in the region of USD 6,000 to USD 15,000. Higher concurrency or larger models push that towards USD 30,000 and beyond. Existing server capacity can sometimes be used for smaller models, which changes the arithmetic considerably.
Can some work stay local and the rest use a hosted model?
Yes, and it is the most common arrangement. Confidential records are processed locally, everything else goes to a hosted model, and the routing rule is written down and auditable. A variation processes documents locally to strip identifying data before anything is sent outside.
Who maintains it once it is running?
Somebody has to, and that is the cost most often left out. Model updates, driver and security patching, monitoring and capacity all become internal responsibilities. Either your own technical team takes it on or it belongs in a support agreement, but it should be named before the hardware is bought.
Related answers
Disclosure, and the only sales pitch on this page
This page is published by a company that sells the thing it describes.
Oxford Crown Technologies builds voice and operations agents for organisations headquartered in London and across the Gulf. We publish these pages as reference material, including market figures that are not ours and that do not always favour us, because a buyer who understands the range negotiates better with everybody, including us. Leadership holding postgraduate degrees from the University of Manchester and the University of Oxford.
Our own builds are fixed fee, live in 14 days, built on accounts in your name, and carry a 28 day money back guarantee. We prove Return on Investment, or you do not pay.
Reviewed 25 July 2026. Market figures change; where this page quotes a range from a third party, the source is named above. Figures described as ours are our published fees or typical results for a first build, not measured results for a named client.