Generative AI Infrastructure Engineer
We are looking for a technical professional to support the development and evolution of the generative infrastructure of Nexum products.
Job Description
We are looking for a technical professional to support the development and evolution of the generative infrastructure of Nexum products. The person will work on local and cloud inference, open-weight model deployment, runtime management, caching, batching, containerization, observability and performance optimization. Tech Stack: Python, Docker, Kubernetes, vLLM, Ollama, llama.cpp, MLX, PostgreSQL, Redis, OpenTelemetry, Llama, Qwen, Mistral, Gemma, DeepSeek, Phi. Key Responsibilities: - Manage and optimize LLM model deployment in production environments - Configure inference runtimes (vLLM, Ollama, llama.cpp, MLX) - Implement caching, batching and load balancing strategies - Containerize and orchestrate AI services with Docker and Kubernetes - Monitor performance with OpenTelemetry and observability tools - Optimize GPU/CPU resource usage for inference This job posting is addressed to both genders, in compliance with laws 903/77 and 125/91 on equal treatment in the workplace and against gender discrimination. We welcome candidates of all ages and nationalities, in accordance with legislative decrees 215/03 and 216/03. Nexum also encourages applications from people with disabilities, in compliance with current regulations.
Additional role details
- Working hours: defined during the selection process for the specific team and project.
- Education requirements: this posting does not specify a mandatory degree; candidates are assessed against the skills and responsibilities listed.
- Experience requirements: this posting does not specify a minimum number of years; relevant, demonstrable experience is assessed against the listed requirements.
What You Bring
AI Infrastructure
Experience with vLLM, Ollama, llama.cpp, MLX and inference runtimes
Containers & Orchestration
Docker, Kubernetes, containerization and model deployment
Observability
OpenTelemetry, monitoring, logging and performance optimization
Database & Caching
PostgreSQL, Redis, caching and batching strategies
How We Work
AI-Assisted Engineering
Approved AI tools accelerate analysis, implementation and documentation; people remain accountable for every released artifact.
Production Perspective
The work connects design, delivery, observability and continuous improvement rather than stopping at isolated prototypes.
Evidence & Traceability
Decisions are supported by tests, metrics, provenance and documented trade-offs.
Cross-Functional Delivery
Cloud, data, AI, security and business requirements are connected around real operating constraints.
Share this open role
Know someone whose skills fit? Share the role or our careers page with your network.

