Generative AI Infrastructure Engineer

We are looking for a technical professional to support the development and evolution of the generative infrastructure of Nexum products.

ItalyEngineering
On-site in ItalyHybrid in ItalyFully remote from Italy

Job Description

We are looking for a technical professional to support the development and evolution of the generative infrastructure of Nexum products. The person will work on local and cloud inference, open-weight model deployment, runtime management, caching, batching, containerization, observability and performance optimization. Tech Stack: Python, Docker, Kubernetes, vLLM, Ollama, llama.cpp, MLX, PostgreSQL, Redis, OpenTelemetry, Llama, Qwen, Mistral, Gemma, DeepSeek, Phi. Key Responsibilities: - Manage and optimize LLM model deployment in production environments - Configure inference runtimes (vLLM, Ollama, llama.cpp, MLX) - Implement caching, batching and load balancing strategies - Containerize and orchestrate AI services with Docker and Kubernetes - Monitor performance with OpenTelemetry and observability tools - Optimize GPU/CPU resource usage for inference This job posting is addressed to both genders, in compliance with laws 903/77 and 125/91 on equal treatment in the workplace and against gender discrimination. We welcome candidates of all ages and nationalities, in accordance with legislative decrees 215/03 and 216/03. Nexum also encourages applications from people with disabilities, in compliance with current regulations.

Additional role details

  • Working hours: defined during the selection process for the specific team and project.
  • Education requirements: this posting does not specify a mandatory degree; candidates are assessed against the skills and responsibilities listed.
  • Experience requirements: this posting does not specify a minimum number of years; relevant, demonstrable experience is assessed against the listed requirements.

What You Bring

  • AI Infrastructure

    Experience with vLLM, Ollama, llama.cpp, MLX and inference runtimes

  • Containers & Orchestration

    Docker, Kubernetes, containerization and model deployment

  • Observability

    OpenTelemetry, monitoring, logging and performance optimization

  • Database & Caching

    PostgreSQL, Redis, caching and batching strategies

How We Work

  • AI-Assisted Engineering

    Approved AI tools accelerate analysis, implementation and documentation; people remain accountable for every released artifact.

  • Production Perspective

    The work connects design, delivery, observability and continuous improvement rather than stopping at isolated prototypes.

  • Evidence & Traceability

    Decisions are supported by tests, metrics, provenance and documented trade-offs.

  • Cross-Functional Delivery

    Cloud, data, AI, security and business requirements are connected around real operating constraints.

Share this open role

Know someone whose skills fit? Share the role or our careers page with your network.

Contact Us

Let's start a great collaboration

Contact Us Illustration