When nothing is allowed to leave the building.

Where public cloud is legally barred, Pexon deploys open-weight models and vector search inside your own data centre. That covers GPU sizing, model deployment and quantisation, an offline retrieval pipeline, and access control bound to your existing directory, so nothing about the system depends on an outside provider.

The starting point

Sometimes the cloud is simply not an option.

Regulation, contractual obligation or a classification decision can rule out sending data to a hosted model. That constraint is real, and it does not have to mean going without the capability.

Sovereign private AI is the position where nothing about the system depends on an outside provider: the model runs on your hardware, the retrieval pipeline runs inside your network, and access control comes from your own directory. The data does not leave, so there is no vendor processing location to review, no renegotiation to schedule and no dependency to explain to an auditor.

The question is no longer whether open-weight models are capable enough — for a large share of enterprise retrieval and extraction work they are, and where they are not, the eval harness shows it before anything ships. The real work is the stack underneath: sizing, deployment, quantisation and operation.

What we build

What running it yourself requires

  1. 01

    Local retrieval and access

    An offline retrieval pipeline with access control bound to your existing directory, so entitlements come from the system of record. The same permission model that governs the source systems travels with the context through retrieval — a user sees exactly what they are entitled to see, no more.

  2. 02

    Model deployment

    Open-weight models deployed and quantised for your hardware, trading precision against throughput deliberately rather than by default. The quantisation choice is validated against your own eval set, so the trade is measured, not assumed.

  3. 03

    Hardware sizing

    GPU selection and capacity planning against your actual workload, so the investment matches the use rather than a benchmark. Sizing runs on rented capacity or existing cards first — the arithmetic comes before the procurement, because hardware bought before the workload is understood is the most expensive mistake in on-premise AI.

Six weeks plus operation.

Sizing comes first: the most expensive mistake in on-premise AI is hardware bought before the workload was understood.

The honest capability question

Open-weight models are not a compromise you must apologise for.

The evaluation that decides whether a sovereign deployment works is the same evaluation that should decide every model choice: your tasks, your documents, your eval set. We do not claim open-weight models beat frontier APIs on every task — we claim the difference is measured in the first weeks, and that for retrieval, extraction and classification work the gap is routinely small enough that the sovereignty requirement settles the decision.

Where the eval shows a task the model cannot carry, the answer is architecture rather than surrender: retrieval quality, prompt structure and routing can move a task from fail to pass without changing the model. The eval harness is the referee, and the engagement hands it over with the system.

The operation question

Who runs it afterwards, and what handover means

A sovereign deployment fails if the team that inherits it cannot operate it. The deliverable therefore includes the runbook, the alerting, the upgrade path and the sizing assumptions — written down, not held in the heads of the engineers who built it. The managed operations tier exists for teams that would rather we kept running it, but it is an option, not a dependency.

That is the sovereignty test in miniature: the system should depend on no outside provider, and that includes us. The architecture, the credentials and the operating knowledge all live on your side, so the relationship can end cleanly at any point without the capability ending with it.

What changes

  • The capability becomes available where a hosted model was never permitted.
  • There is no outside dependency to review, renegotiate or explain.
  • Your data classification stops being the reason the project cannot start.
  • Operation stays with your team: runbooks and documentation are part of the deliverable, with a managed operations tier only if you want it.

Sovereignty questions

Are open-weight models good enough?

For a large share of enterprise retrieval and extraction work, yes — and where they are not, that shows up in the eval harness before anything ships. The honest answer depends on the task, which is what the first weeks establish.

Who operates it afterwards?

Your team, by design. We hand over runbooks and documentation, and the managed operations tier exists only if you would rather we kept it running.

Next step

Not a sales call. An architecture call.

Thirty minutes with the architect who would actually run the engagement.