The data cannot be reached. That is the actual problem.
The data foundation is the layer that makes enterprise systems reachable for AI. Pexon extracts, cleans, masks and indexes data from SAP, MES, PLM, SharePoint and legacy databases, then carries your real permission model through end to end so the resulting system can pass a security review.
Reachability first, model second.
Every AI project in an industrial company hits the same wall: the model is ready and the use case is clear, but the knowledge it needs lives in systems that were never designed to be queried — thirty-year-old schemas, proprietary interfaces, data spread across SAP, MES and PLM with no single view. The data foundation is the layer that removes that wall once, so every application after it starts from a position where the data is actually reachable.
This is not an integration project in the usual sense. It is extraction, cleaning, masking and indexing of production data with your real permission model carried through end to end — the difference between a system that can go to production and one that stays a demo.
A prompt cannot reach into a thirty-year-old schema.
Retrieval is only as good as the data it can reach, and in industrial companies the authoritative data sits behind interfaces designed for people, not for software. Making it reachable takes engineers who read your data model, understand which fields are authoritative — which plant, which material master, which document version — and can express your entitlements in the retrieval layer.
That is why the foundation comes before the brain: the company brain sits on top of this layer, and every capability it offers — retrieval, agents, monitoring — inherits the reachability and the permission model built here. Skipping the foundation and buying a generic retrieval platform means re-integrating the same sources twice, once per project, forever.
Three stages, one contract with the data
The foundation is a pipeline with a fixed shape. Each stage has a property that cannot be skipped without turning the whole layer into a demo.
| Stage | What it guarantees |
|---|---|
| Extract | Read-only access to SAP, MES, PLM, SharePoint and legacy SQL — against replicas or extracts wherever your operations team prefers, with no write path by default. |
| Clean and mask | Field-level cleaning, PII masking before anything is indexed, and masking rules that are part of the build rather than an afterthought. |
| Index with permissions | The data is indexed together with your real permission model, so the retrieval layer can answer the question "who is allowed to see this record" at query time. |
Where the reachability is built, one source at a time
The engagements on this page are the source-system paths we build first. Each one is a specific extraction problem with its own schema, its own authoritative fields and its own permission model — which is exactly why they are built by engineers who read the data model rather than by connectors configured from a dropdown.
- SAP data layer — extraction from the SAP estate with field-level authority and masking, described on the SAP data layer page.
- MES production data — machine and process data from the manufacturing execution layer, described on the MES production data page.
- Legacy code modernisation — reaching into the systems nobody wants to touch but everybody depends on, described on the legacy code modernisation page.
Read-only by default. Masked before indexing. Audited from the first sprint.
Three constraints shape every foundation build. First, no write access to production systems: we connect read-only by default and work against replicas or extracts wherever your operations team prefers, with write paths added later, deliberately, and only where a use case requires them. Second, personal data is masked in the pipeline before anything is indexed — the masking rules are part of the build, not an afterthought. Third, audit logging is on from the first sprint, so the layer can show its own history when the security review asks.
These are the constraints that let the resulting system pass a security review instead of failing one. A data layer that cannot answer where a record came from, who can see it and what was masked is not a foundation — it is a liability the next project inherits.
Source systems we start from
Legacy Code Modernisation — Document the Rules First
Business rules extracted from legacy code and mainframe systems, documented and covered by generated tests, so modernisation becomes possible.
MES Production Data for AI — Root Cause in Hours
Historised machine and batch context joined to quality results, so recurring faults become visible and root-cause analysis stops taking weeks.
SAP Data Layer for AI — One Answer, Traceable to the Row
A governed semantic layer over SAP master and transactional data, with metric definitions fixed at source so orders, margins and stock have one answer.
Common questions
Do you need write access to our production systems?
No. We connect read-only by default and work against replicas or extracts wherever your operations team prefers. Write paths are added later, deliberately, and only where a use case requires them.
What happens to personal data?
PII is masked in the pipeline before anything is indexed, and the masking rules are part of the build rather than an afterthought. Audit logging is on from the first sprint.
Not a sales call. An architecture call.
Thirty minutes with the architect who would actually run the engagement.
