Open-weight model hosting: the checkpoint runs where the data already is.

An open-weight model is a language model whose weights you download and run yourself. Pexon hosts Nemotron, Qwen, GLM and DeepSeek on GPU servers in the EU, behind a LiteLLM gateway, so inference never leaves the tenant. The two-week Readiness Blueprint sizes the open-weight model and the hardware before anyone orders cards.

The starting point

The API is the experiment. The open-weight model is the production constraint.

Industrial companies between EUR 50m and 2bn revenue usually meet the wall the same week: a classification rule, a customer contract, or a token bill that no longer looks like a pilot. An open-weight model is the checkpoint you can download and run on GPUs you control. Nemotron, Qwen, GLM and DeepSeek are the families we put behind a LiteLLM gateway so the application never learns a vendor URL.

The model is free to download. The deployment is the product.

What we build

Three layers, in this order

01

The model shortlist against your eval

Nemotron when western provenance and NVIDIA's stack matter, Qwen when the size ladder on fewer cards wins, GLM or DeepSeek when the eval already says so. Ultra is a 550B/55B cluster SKU with a published 8× B200 or 16× H100 floor. We do not start there unless Super-class checkpoints have already lost.

02

The serving stack on your hardware

vLLM or NIM on the cards the sizing arithmetic named, with quantisation treated as a measured trade. The gateway in front is LiteLLM: one API, budgets, audit logs. Prompts and outputs stay in the tenant.

03

The runbook you can operate

Upgrade path, eval harness, and the honest split: tasks that stay on the open-weight model, and the slice that still needs a frontier API. We write that split down. We do not hide it behind a single logo.

What changes

  • Inference location becomes a configuration you can show an auditor, not a vendor region on a slide.
  • Token spend on the bulk of traffic becomes a GPU and power line, which finance can plan, instead of an invoice nobody owns.
  • Model swaps happen in the gateway. The application keeps one base URL when Nemotron, Qwen or a frontier API takes a given route.
  • Hardware is ordered after the eval, so you do not buy sixteen H100s to discover a 32B Qwen already passed.
  • The limit is written down: if the open-weight model fails a named task, that task stays on a frontier API behind the same gateway.

Sources: NVIDIA Nemotron (open weights, data, recipes) · Nemotron 3 Ultra model card (550B/55B, GPU floor, OpenMDW 1.1). Read 30 August 2026. Vendor documentation changes; verify against the current release.

Start with the two-week Readiness Blueprint.

€4,900 two weeks, fixed price

We run your tasks against the candidate open-weight models, size the GPUs, and hand you a gatewayed serving plan with every assumption written down. The document is yours to keep, whether Pexon builds the stack or your existing integrator does.

All prices are net and exclude VAT.

Open-weight hosting questions

What does open-weight model hosting cost?

The two-week Readiness Blueprint is fixed at €4,900 net. It returns the model shortlist, the GPU count and a gatewayed serving plan. A production use-case pilot starts from €15,000. Hardware is quoted after the measurement, never before.

What is an open-weight model?

An open-weight model is a language model whose parameters you can download and run on hardware you control. Nemotron, Qwen, GLM and DeepSeek are the families we put on that path. The licence still has to be read; open weights are not a blank cheque, and they are not the same thing as a fully open dataset in every case.

Do we need Nemotron 3 Ultra to start?

No. Ultra is a 550-billion-parameter cluster SKU with a published floor of 8× B200-class or 16× H100 GPUs. Most industrial retrieval and assistant workloads start on a smaller Nemotron tier or on Qwen. The blueprint runs the eval before the card count is frozen.

Where are the limits of this offer?

We do not pretend every task belongs on an open-weight model. If the eval fails the self-hosted candidate, we say so and keep a frontier API behind the same gateway for that slice. Sovereignty and cost are the reasons to host; they are not a reason to lie about quality.

Next step

Not a sales call. An architecture call.

Thirty minutes with the architect who would actually run the engagement.