What we build.
Pexon builds AI in six clusters: private AI on hardware you own, Claude rolled out across the enterprise, reaching the data locked in SAP, MES and PLM, a company brain on top of it, governance over what models may see, and inference cost control. Every engagement starts with one source taken end to end into production.
Production-Ready AI Engineering.
Most AI providers position as consulting-plus-development or as an agency. The category that stays empty is the one where GenAI systems actually run in production — and somebody operates them. That is the category this page is built around.
Production means the model is not the deliverable: monitoring, evals, observability and CI/CD for models are the core of the engagement, not a side note. We take the system into production in your tenant, and we operate it from there — that is the part agencies do not carry.
AI in production needs six things, and we build all six.
Every AI programme that reaches production in an industrial company needs the same six capabilities, and the order they are built in decides whether the programme survives: reach the data, answer from it, run the models you are allowed to run, govern what they may see, and control what they cost. The six clusters on this page are exactly those six jobs.
You do not have to buy the programme. Each cluster begins as a fixed-price two-week blueprint, and the blueprint is yours to keep whoever ends up building the system — the architecture plan is the deliverable, not a sales step towards one.
The order is deliberate.
Most AI programmes start at the use case and discover the data problem in month four. We start at the data layer, because reachability and permissions are what decide whether anything reaches production at all. A use case on unreachable data is a demo; a data layer without a use case is a foundation waiting for one.
The same logic runs through the sequence below: the data foundation comes before the company brain, because the brain answers from the foundation. Governance and cost come last not because they matter less, but because they are the same architecture — the gateway that logs is the gateway that routes — so they are built into the earlier clusters rather than bolted on afterwards.
Where to go, depending on what is blocking you
Each cluster below is a hub with its own pages. If you are deciding where to start, pick the one that names your current blocker — each entry links the engagement we would start with.
- Data foundation. Reach the data that is actually locked away: SAP, MES, PLM and legacy systems extracted, masked and permission-aware, so AI can go to production instead of staying a demo. Start with the SAP data layer.
- Company brain. Answer from what the company already knows: grounded retrieval over your own documents and records, permission-aware and evaluated before release. Start with the enterprise RAG platform.
- Claude across the enterprise. Run Claude the way an industrial company has to run it: a gateway for control, MCP for integration, EU hosting and GDPR where the data demands it. Start with the LLM gateway for Claude.
- Private AI. Own the model: open-weight models on hardware you control, sized to your workload instead of a benchmark, served and operated by your own team. Start with GPU inference sizing.
- AI governance. Govern what models may see: one gateway for all AI traffic, PII interception, audit-grade logging, and sovereign deployment where the law requires it. Start with the shadow AI lockdown.
- AI FinOps. Control the bill: semantic caching, model routing and per-department attribution, so the escalating inference spend becomes a predictable line item. Start with LLM cost optimisation.
One source, end to end, into production.
The rule for choosing a first engagement: pick the source system that blocks the most people. We take one source end to end — extraction, masking, indexing, entitlements, evaluation — rather than sampling four in parallel, because a thread that reaches production teaches you more than four pilots that do not.
Every cluster can be entered the same way: a two-week blueprint that maps your landscape, assesses the permission model and hands you a costed architecture plan. Stop after it if you want to — you will have paid for a plan and received one.
Six clusters
AI Server — Buy or Build the Right Hardware for LLM Inference
AI server decisions: GPU sizing, Ollama vs vLLM serving, cost per query, and when to buy, rent or run in the cloud.
Claude for Enterprise: Three Things That Stall a Rollout
Most Claude enterprise rollouts do not fail on the model. They stall on three things: the permission model, the tenant boundary and the cost model.
Private AI — Open-Weight Models on Hardware You Own
Open-weight models on hardware you own: GPU sizing, quantisation, a vLLM serving layer and access control bound to your own directory.
AI FinOps — Get the Model Bill Under Control
Semantic caching, model routing by query complexity and per-department usage visibility, so an escalating inference bill becomes a predictable one.
AI Governance — Compliance as Architecture, Not Paperwork
One gateway for AI traffic, PII interception, audit-grade logging and EU AI Act readiness, built into the system rather than documented after it.
Company Brain — AI Systems Built on Your Own Knowledge
Retrieval, agents and monitored models that answer from your own data, respect who is allowed to see what, and run inside your tenant.
Data Foundation — Make Enterprise Data Reachable for AI
SAP, MES, PLM, SharePoint and legacy SQL, extracted, cleaned, masked and permission-aware, so an AI system can go to production instead of staying a demo.
Common questions
Where should we start if we have several candidate use cases?
With the source system that blocks the most people. We take one source end to end rather than sampling four in parallel, because a thread that reaches production teaches you more than four pilots that do not.
Do we have to buy the whole programme?
No. Each cluster begins as a fixed-price two-week blueprint. You can stop after it and keep the architecture plan, whoever ends up building the system.
Not a sales call. An architecture call.
Thirty minutes with the architect who would actually run the engagement.
