Generative AI Infrastructure Engineer
We are looking for a technical professional to support the development and evolution of the generative infrastructure of Nexum products.
Job Description
We are looking for a technical professional to support the development and evolution of the generative infrastructure of Nexum products. The person will work on local and cloud inference, open-weight model deployment, runtime management, caching, batching, containerization, observability and performance optimization. Tech Stack: Python, Docker, Kubernetes, vLLM, Ollama, llama.cpp, MLX, PostgreSQL, Redis, OpenTelemetry, Llama, Qwen, Mistral, Gemma, DeepSeek, Phi. Key Responsibilities: - Manage and optimize LLM model deployment in production environments - Configure inference runtimes (vLLM, Ollama, llama.cpp, MLX) - Implement caching, batching and load balancing strategies - Containerize and orchestrate AI services with Docker and Kubernetes - Monitor performance with OpenTelemetry and observability tools - Optimize GPU/CPU resource usage for inference This job posting is addressed to both genders, in compliance with laws 903/77 and 125/91 on equal treatment in the workplace and against gender discrimination. We welcome candidates of all ages and nationalities, in accordance with legislative decrees 215/03 and 216/03. Nexum also encourages applications from people with disabilities, in compliance with current regulations.
What You Bring
AI Infrastructure
Experience with vLLM, Ollama, llama.cpp, MLX and inference runtimes
Containers & Orchestration
Docker, Kubernetes, containerization and model deployment
Observability
OpenTelemetry, monitoring, logging and performance optimization
Database & Caching
PostgreSQL, Redis, caching and batching strategies
How We Work
AI-Assisted Engineering
Approved AI tools accelerate analysis, implementation and documentation; people remain accountable for every released artifact.
Production Perspective
The work connects design, delivery, observability and continuous improvement rather than stopping at isolated prototypes.
Evidence & Traceability
Decisions are supported by tests, metrics, provenance and documented trade-offs.
Cross-Functional Delivery
Cloud, data, AI, security and business requirements are connected around real operating constraints.

