Generative AI Infrastructure Engineer

We are looking for a technical professional to support the development and evolution of the generative infrastructure of Nexum products.

ItalyEngineering

Job Description

We are looking for a technical professional to support the development and evolution of the generative infrastructure of Nexum products. The person will work on local and cloud inference, open-weight model deployment, runtime management, caching, batching, containerization, observability and performance optimization. Tech Stack: Python, Docker, Kubernetes, vLLM, Ollama, llama.cpp, MLX, PostgreSQL, Redis, OpenTelemetry, Llama, Qwen, Mistral, Gemma, DeepSeek, Phi. Key Responsibilities: - Manage and optimize LLM model deployment in production environments - Configure inference runtimes (vLLM, Ollama, llama.cpp, MLX) - Implement caching, batching and load balancing strategies - Containerize and orchestrate AI services with Docker and Kubernetes - Monitor performance with OpenTelemetry and observability tools - Optimize GPU/CPU resource usage for inference This job posting is addressed to both genders, in compliance with laws 903/77 and 125/91 on equal treatment in the workplace and against gender discrimination. We welcome candidates of all ages and nationalities, in accordance with legislative decrees 215/03 and 216/03. Nexum also encourages applications from people with disabilities, in compliance with current regulations.

What You Bring

  • AI Infrastructure

    Experience with vLLM, Ollama, llama.cpp, MLX and inference runtimes

  • Containers & Orchestration

    Docker, Kubernetes, containerization and model deployment

  • Observability

    OpenTelemetry, monitoring, logging and performance optimization

  • Database & Caching

    PostgreSQL, Redis, caching and batching strategies

How We Work

  • AI-Assisted Engineering

    Approved AI tools accelerate analysis, implementation and documentation; people remain accountable for every released artifact.

  • Production Perspective

    The work connects design, delivery, observability and continuous improvement rather than stopping at isolated prototypes.

  • Evidence & Traceability

    Decisions are supported by tests, metrics, provenance and documented trade-offs.

  • Cross-Functional Delivery

    Cloud, data, AI, security and business requirements are connected around real operating constraints.

Contact Us

Let's start a great collaboration

Contact Us Illustration