Product knowledge in seconds, not half an hour of searching

A German transformer manufacturer with 4,500 employees cut research per technical enquiry from up to 30 minutes to seconds. Pexon built an LLM-independent assistant on Open WebUI inside the company's own Azure tenant, indexing SharePoint and Confluence behind Entra ID. Licence cost is zero; infrastructure runs at EUR 3 to 5 per user per month.

Outcome
−90%
Research time per technical enquiry
to seconds
From up to 30 minutes
EUR 0
Recurring licence cost
EUR 3–5
Infrastructure per user, per month

The engagement at a glance

Written for the head of technical support or service operations at a manufacturer whose product knowledge is spread across SharePoint, Confluence and SAP, and who has watched experienced people spend their afternoons finding documents rather than answering customers.

CustomerGerman manufacturer of transformer technology, more than 4,500 employees, around EUR 1 billion revenue
IndustryEnergy technology, transformers, industrial manufacturing
UsersTechnical support, maintenance and internal sales
CloudMicrosoft Azure, the customer's own tenant
StackOpen WebUI, Entra ID, SharePoint, Confluence, vector search pipeline, product metadata filtering, image and diagram processing
Timeframe2025 to 2026
ComplianceGDPR, industrial data protection, no document content leaves the customer tenant
StatusProductive in daily operations, not a pilot

Thirty minutes to answer a question the company already knew the answer to

The knowledge existed. Finding it was the job. Staff in technical support, maintenance and internal sales needed product-specific detail from operating manuals, technical wikis and internal documentation many times a day, and that material was spread across SharePoint, Confluence and SAP with no common search across them. The average enquiry cost around thirty minutes of manual searching before anyone could start answering it.

Thirty minutes is not the whole cost. The company sells a broad portfolio of actively distributed products, so the search is not one lookup but a sequence: find the right source system, then the right document, then the right passage, then check it applies to the specific product variant in front of you. Every step is a chance to stop at a plausible-looking answer from the wrong revision.

Underneath that sat a more expensive problem. Product knowledge was tied to individual experts. When one of them was on holiday or left, the knowledge went with them: deputising was difficult, onboarding was slow, and the organisation carried a quiet dependency on a handful of people remembering things. That is a risk no document management project had ever fixed, because the documents were never the bottleneck. The retrieval was.

What was built: one assistant, the company's own tenant, no licence per seat

Pexon built a data-protection-compliant AI platform that makes SharePoint, Confluence and internal documentation answerable through a single chat interface, running entirely inside the customer's own Azure environment. No document content is transferred outside it. That was a project condition, not a preference, because the indexed corpus includes technical documentation and construction data that has no business leaving the company.

The retrieval layer does more than embed documents and hope. Product metadata filtering constrains the search to the variant actually being asked about, which is the difference between an assistant that is useful in a product company and one that confidently quotes the wrong generation of a device. Images and diagrams are processed too, because in transformer technology a meaningful share of the answer lives in a schematic rather than a paragraph.

Authentication and authorisation run through Entra ID single sign-on. The assistant answers from the documents the person asking is already entitled to read, and entitlement is resolved at retrieval time rather than left to the model to respect. This is the part that decides whether a system like this can go into production at all: an assistant that can quote a document the user could not have opened is a data-protection incident with a chat interface.

Trade-off: The obvious alternative was Microsoft 365 Copilot, and it would have been faster to buy than to build. It was rejected on architecture and on cost. Copilot ties the company to one model provider and prices per seat, which means the bill scales with adoption in exactly the way that makes internal rollouts stall. The LLM-independent route on Open WebUI cost more engineering up front and gives the company the right to change model without changing platform, at zero recurring licence cost. The price of that freedom is that the connectors, the permission synchronisation and the metadata filtering are yours to build and yours to maintain.

Why LLM independence was worth paying for

Model independence sounds like an architecture preference until a model is deprecated. The platform treats the language model as a swappable component behind a stable interface, so a change of provider or a change of model version is a configuration decision rather than a migration project. In a market where model generations turn over faster than enterprise procurement cycles, that is the difference between a platform and a purchase.

It also changes the cost conversation permanently. Recurring licence cost is zero. Infrastructure runs at roughly EUR 3 to 5 per user per month, which is materially below comparable SaaS assistants and below per-seat assistant licensing, and the gap widens with every user added rather than narrowing. Cost per user going down as adoption goes up is the opposite of the usual dynamic, and it is why this platform could be opened to the whole support organisation instead of a pilot group.

  • Model as a replaceable component, not a dependency written into the platform
  • EUR 0 recurring licence cost across the user base
  • Around EUR 3 to 5 per user per month in infrastructure
  • Reusable foundation: further use cases in HR support, customer support and onboarding were built on the same platform

What actually changed for the people using it

Research time per technical enquiry fell by more than 90%. Questions that had cost up to thirty minutes of manual searching are answered in seconds, and the answer arrives with its source rather than as an assertion.

The deputising problem moved from unsolved to solved. For the first time, standing in for an absent expert became systematically possible without the knowledge loss that used to come with it, because the knowledge is now in a system that answers rather than in a person who is away. That was the outcome the company had been trying to buy for years through documentation initiatives, and it turned out to be a retrieval problem all along.

The system is in productive daily use. It is not a pilot, not a proof of concept and not a demo that impressed a steering committee, which is the distinction that matters when you are deciding whether an approach is real.

Before and after

DimensionBeforeAfterEffect
Time per technical enquiryUp to 30 minutes of manual searchSecondsMore than 90% reduction
Where knowledge livedSharePoint, Confluence and SAP, searched separatelyOne assistant across all of themFind, then answer becomes just answer
Cover for absent expertsDifficult, with knowledge lossSystematically possibleDependency on individuals reduced
Cost modelPer-seat licensing under evaluationEUR 0 licence, EUR 3–5 infrastructure per userCost per user falls as adoption rises

Questions buyers ask about this

What does an internal AI product support assistant cost per user?

In this deployment, EUR 3 to 5 per user per month in infrastructure, with no recurring licence cost, because the architecture is LLM-independent and runs on Open WebUI inside the company's own Azure tenant. That is materially below per-seat assistant licensing, and the gap widens with every additional user.

How long does it take to build one?

This system went from start to productive daily use across 2025 and 2026, including the connectors to SharePoint and Confluence and the permission synchronisation through Entra ID. A first working assistant over one document source is a two-week blueprint; the schedule is set by how many source systems must be indexed and how messy their permissions are.

Does an AI assistant like this leak internal documents to a model provider?

Not in this architecture. The whole system runs inside the customer's own Azure environment and no document content is transferred outside it. That was a condition of the project rather than an optimisation, because the indexed material includes technical documentation and construction data.

How does it handle who is allowed to see which document?

Access is resolved through Entra ID single sign-on, so the assistant answers from the documents the person asking is already entitled to read. Retrieval is filtered by that entitlement rather than by the model being asked politely not to mention things, which is the only version of this that survives an audit.

Does this transfer to industries other than energy technology?

The pattern transfers wherever product knowledge sits in SharePoint, Confluence or a wiki and support staff search it by hand. The engineering that has to be redone per company is the connector layer and the permission model, not the assistant itself. The same platform in this case later carried further use cases in HR and customer support.

Next step

Your product knowledge is probably in the same three systems

If support staff at your company are spending half an hour finding what the company already knows, the two-week blueprint answers the only questions that matter before you commit: which source systems can actually be indexed, how the permission model resolves, and what it costs per user at your headcount.