Open role · Cluj-Napoca · Full time

DevOps Engineer

Make private AI a platform people can rely on.

  • Kubernetes
  • Cloud
  • Inference

The role

Make private AI a platform people can rely on.

Pexon Romania is hiring a DevOps and cloud native engineer in Cluj-Napoca to run the platforms its AI work is delivered on: Kubernetes clusters, GPU nodes serving open-weight models, and the gateway that meters and logs every call. Jobs are based at the Cluj office and the work happens in customer tenants.

The work

The platform you will own, layer by layer.

03 / Gateway

One front door for models

Teams need access to models without losing control of usage and cost.

Your part
Configure routing, quotas, caching and audit logging with gateways such as LiteLLM.
The outcome
Model access that can be governed and explained.
  • LiteLLM
  • Quotas
  • Audit logs
Explore model access and governance

02 / Serving

Inference that holds up under load

GPU memory, batching and concurrency change capacity planning.

Your part
Deploy serving tools such as vLLM and size workloads before hardware is committed.
The outcome
A serving layer designed around real demand.
  • vLLM
  • GPU sizing
  • Batching
Explore GPU inference sizing

01 / Foundation

Kubernetes in the customer's tenant

The platform must fit existing identity, networking and change processes.

Your part
Build node pools, storage and delivery pipelines within those constraints.
The outcome
A platform with observable behaviour and a clear operating handover.
  • Kubernetes
  • Networking
  • CI/CD
Explore private AI platforms

Across every layer: CI/CD, observability, runbooks and a handover the customer’s team can use.

What you bring

A production mindset. Solid foundations.

  • Kubernetes in production

    You have debugged real scheduling or runtime problems and can explain how you approached them.

  • One major cloud in depth

    AWS, Azure or Google Cloud, including identity, networking and quotas.

  • Containers and CI/CD

    Docker and a delivery pipeline you have built with GitLab CI, GitHub Actions or Jenkins.

  • Code for automation

    Python or Go for tooling, operators and integration work.

  • Technical English

    Comfort working through platform decisions with customer engineers.

Room to grow

What you can learn here.

  • GPU inference and model serving
  • How batching and quantisation affect capacity
  • Token costs alongside latency and reliability

Prior GPU experience and German are welcome, but neither is a requirement. You operate models here; the role does not require a machine-learning research background.

A person behind the process

Meet Laura Picovici.

HR Business Partner · Recruiting

Laura runs recruiting for the Cluj hub. Your application reaches her first, she answers your questions about the process and the range, and she brings in the engineers you would actually work with. Later in the process you also meet Noel Dinger, who runs the company.

Meet the people behind Pexon

The conversation

A clear path from hello to a decision.

  1. Introduce yourself

    Share something you have built or operated. Your message reaches Laura Picovici, who runs recruiting for the hub.

  2. Talk engineering

    A 30–60 minute first conversation about your work and what you want next. Ask for the salary range by email beforehand.

  3. Work through a problem

    A practical session together, inside working hours. No weekend take-home assignment.

  4. Meet the hub

    Talk with Noel Dinger about the company and what you would own.

  5. Get a decision

    A yes or a no, with a reason. After a technical session, feedback addresses the work discussed.

The practical questions.

Where is the role based, and what is the office schedule?

The role is based at the Cluj-Napoca office. The team is built around people who can work there together. Discuss the current office schedule in the first conversation; this is not a remote-first role.

Can I ask for the salary range before a call?

Yes. Ask in your application email and the range will be shared in the reply, before a technical session. The email preparation form includes an optional request for the range so you can raise it from the start.

Do I need to have served models on GPUs before?

No, and few people have. What is needed is solid Kubernetes and enough curiosity about what makes inference different from a web workload: memory that is reserved rather than requested, batching that changes throughput by an order of magnitude, and a cold start measured in minutes. The rest is learnable here and is being learned here.

Whose infrastructure am I working in?

The customer's, in almost every case. That is what forward deployed means: their cloud subscription, their cluster, their security review and their change window. It is the part of the job people underestimate, because a platform you do not own is a platform where the fastest fix is often not available to you.

Is this an on-call role?

What the on-call arrangement looks like depends on the engagement, and the honest answer is that it is negotiated per customer rather than fixed by a company policy this page could quote. Ask about it in the first conversation; the answer you get will be about the specific engagements currently running.

How is this different from a platform role at a product company?

You will run several platforms rather than one, in environments you did not design, for teams that are not yours. The variety is the appeal and it is also the cost: you rarely get the luxury of a two-year plan to fix a foundation, and you have to be able to leave a system in a state someone else can operate.

Your next move

Start with something you have built.

Tell us about a system, a difficult problem, or something you made better. A short introduction is enough to start the conversation.

Apply for this role

Prepare an email. Review it. Send when you are ready.

Another direction

Prefer building the software?

View all open roles
Engineering · Cluj-Napoca

Software Engineer

Build the systems around the model.

  • Python / Java
  • APIs
  • Retrieval
Full time · Office-basedView role