Systems and data

Can AI run on your own servers, so data never leaves?

Yes. Open-weight models can be run on hardware your organisation owns, inside your own network, so that no document, recording or record is ever transmitted to a third party. It costs more up front than calling a hosted model, and the very best models are not available this way, but for extraction, classification, summarising and routine drafting the gap is small enough that most regulated workloads run perfectly well.

Last reviewed . Written and maintained by Oxford Crown Technologies.

The three deployment models

Almost every conversation about data residency collapses into one of these three, and a great deal of confusion comes from suppliers using the word private for all of them.

ModelWhere the data goesTypical cost shapeSatisfies
Hosted model interfaceTo the provider, processed in their region of choice unless restrictedPer item. No capital outlayMost commercial confidentiality requirements, with a data processing agreement
Dedicated cloud region or private networkStays in a named jurisdiction, on infrastructure the provider operatesPer item, plus a platform or commitment feeResidency clauses that specify a country but not ownership of the hardware
On your own hardwareNever leaves your networkCapital cost up front, then power and maintenanceClauses that require the data to remain on infrastructure you control, and air-gapped environments
The middle option is the one most organisations actually need and the one most often skipped in favour of the extremes.

What running it yourself actually involves

An open-weight model is a file. It is downloaded, placed on a server with a suitable graphics processor, and served through a local interface that the rest of your systems call in the same way they would call a hosted provider. From the point of view of the applications sitting on top, almost nothing changes. What changes is that you now own the uptime, the patching, the capacity and the upgrade path.

When residency genuinely forces it

This is worth separating from preference, because on-premise costs more and should be chosen for a reason that can be written down.

  1. A contract says so

    Client agreements, particularly in defence, government supply chains and parts of financial services, frequently require that data is not transmitted to a third party processor. This is the most common genuine trigger, and it is a document you can point at.

  2. The data cannot lawfully leave the country

    Several jurisdictions, including a number across the Gulf, place residency requirements on personal, health or government-related records. A dedicated region sometimes satisfies this, and sometimes the requirement is explicitly about who operates the infrastructure.

  3. The material is confidential in a way that survives a contract

    Unreleased financial results, live legal matters, merger and acquisition work and clinical records are cases where organisations reasonably decline to rely on a processing agreement, however well drafted.

  4. There is no outbound connectivity

    Industrial sites, secure facilities and some field operations run without a route to the public internet at all. This is the one case where the decision requires no debate.

The cost, honestly

On-premise trades a per item charge for a capital cost and an ongoing obligation. The crossover depends almost entirely on volume.

Hosted modelOn your own hardware
Up frontNothingUSD 6,000 to 15,000 for a single professional GPU server, and USD 30,000 and beyond for higher concurrency
Per itemFractions of a cent to a few cents, depending on model and lengthEffectively nothing once the hardware is bought
OngoingUsage onlyPower, rack space or floor space, patching, monitoring and eventual replacement
Who fixes it at 3amThe providerYou, or whoever holds your support agreement
Model upgradesAutomatic, sometimes whether you wanted them or notA deliberate project each time, which is an advantage for stability and a cost in effort
Best suited toLow and medium volume, and anything needing the strongest reasoningHigh steady volume, and anything where the data cannot leave
Hardware figures are approximate market ranges for mid-2026 and move quickly. Existing server capacity can sometimes run smaller models, which changes the arithmetic considerably.

The commercial argument that is often missed: an on-premise build is a fixed cost that does not grow with the size of your team. Per seat licensing for commercial assistants scales with headcount, so at forty or fifty users the comparison starts to look very different from the way it looks at five. That is one of the few cases where privacy and cost point the same way.

The capability trade-off

Open-weight models are not equivalent to the strongest hosted models and it is unhelpful to pretend otherwise. The gap is uneven, which is the useful part.

  • Extraction, classification, routing, summarising and routine drafting run close to parity. These are also the great majority of business automation workloads.
  • Long multi-step reasoning, difficult code and unusual analytical work are where the leading hosted models remain measurably ahead.
  • Speed on your own hardware is a function of what you bought. A single professional GPU serving a small team is comfortable, and the same hardware serving two hundred users is not.
  • The gap has narrowed every year and there is no reason to assume it stops. A workload that is marginal today is a reasonable candidate for review in twelve months.
  • The only test that matters is your own. Run fifty of your real, hardest cases through both and compare the output, rather than relying on any published claim, ours included.

The middle path most organisations end up at

Few builds are entirely one or the other. Two patterns cover most of what we see in practice.

  1. Route by sensitivity

    Confidential records are processed locally and everything else goes to a hosted model. The routing rule is written down, testable and auditable, which matters more than where any individual item ends up.

  2. Redact, then send

    A local model strips names, account numbers and other identifying data, and only the redacted text leaves the network. This gives most of the capability of a frontier model with a much smaller disclosure. It requires the redaction step itself to be tested properly, because a leak here is silent.

Six questions before committing to on-premise

  • Which specific clause, regulation or constraint requires it, in writing.
  • Whether a dedicated region would satisfy that clause, since it usually costs far less.
  • How many people and how many items a day the system has to serve, because that determines the hardware and therefore most of the cost.
  • Who patches, monitors and upgrades it, named before the hardware is ordered.
  • What the fallback is when the server is down, because a local deployment has no provider status page and no automatic failover unless somebody built one.
  • How the model will be evaluated at upgrade time, so that a newer version cannot quietly regress on the cases you depend on.

Our own position, marked as ours

We build both, and we do not treat on-premise as the default. Where a contract, a regulation or an absent network connection forces it, the models run on hardware you own, nothing leaves the premises, and the build is a fixed cost rather than a licence that grows with your team. Where it does not, we say so, because recommending hardware nobody needed is an expensive way to be cautious. That figure of one fixed build instead of a per seat licence is typical of a first build rather than a measured result for a named client.

Frequently asked questions

Is a private cloud region the same as on-premise?

No, although it is often good enough. A dedicated region or a virtual private cloud keeps data inside a jurisdiction and outside the model provider's training pipeline, but the infrastructure still belongs to somebody else. On-premise means the hardware is in your building and the network boundary is yours. Only the second satisfies the strictest contractual clauses.

How far behind are open-weight models?

On extraction, classification, summarising and routine drafting they are close to parity and the difference rarely shows in production output. On long multi-step reasoning and difficult code the leading hosted models remain measurably ahead. The practical approach is to test your own hardest cases against both rather than trust either claim.

What hardware is actually needed?

For a single team's workload, one server with a professional GPU is usually sufficient, in the region of USD 6,000 to USD 15,000. Higher concurrency or larger models push that towards USD 30,000 and beyond. Existing server capacity can sometimes be used for smaller models, which changes the arithmetic considerably.

Can some work stay local and the rest use a hosted model?

Yes, and it is the most common arrangement. Confidential records are processed locally, everything else goes to a hosted model, and the routing rule is written down and auditable. A variation processes documents locally to strip identifying data before anything is sent outside.

Who maintains it once it is running?

Somebody has to, and that is the cost most often left out. Model updates, driver and security patching, monitoring and capacity all become internal responsibilities. Either your own technical team takes it on or it belongs in a support agreement, but it should be named before the hardware is bought.

Disclosure, and the only sales pitch on this page

This page is published by a company that sells the thing it describes.

Oxford Crown Technologies builds voice and operations agents for organisations headquartered in London and across the Gulf. We publish these pages as reference material, including market figures that are not ours and that do not always favour us, because a buyer who understands the range negotiates better with everybody, including us. Leadership holding postgraduate degrees from the University of Manchester and the University of Oxford.

Our own builds are fixed fee, live in 14 days, built on accounts in your name, and carry a 28 day money back guarantee. We prove Return on Investment, or you do not pay.

Reviewed 25 July 2026. Market figures change; where this page quotes a range from a third party, the source is named above. Figures described as ours are our published fees or typical results for a first build, not measured results for a named client.