ML Inference Platform Engineer
Build the infrastructure that serves models reliably and makes latency, throughput and operating cost visible.
- Location
- Kyiv / remote
- Experience
- Senior
- Working arrangement
- Arrangement to be discussed
Requirements
- Experience operating backend or ML infrastructure under production load.
- Strong Linux and networking fundamentals, with Python, Go, Rust or C++.
- Ability to compare serving approaches using realistic workloads and cost constraints.
Responsibilities
- Design model-serving paths, batching, routing and rollout mechanisms.
- Profile GPU utilisation, memory pressure, tail latency and failure recovery.
- Build load tests, observability and repeatable deployment procedures.
Useful evidence
- An infrastructure project with measured performance and operational lessons.
- GPU profiling, distributed inference and custom kernels are useful experience.
Working terms
- Startup culture, a goal-oriented team, and a research mindset
- The opportunity to apply your engineering skills to tools and systems for fellow engineers and help shape the future of AI
- Latest-generation MacBook Pro
- An in-house GPU cluster for training and experimentation
- 20 working days of annual leave
- English courses, educational events, and conferences
- Medical insurance
Tools & systems
NVIDIA DynamovLLMTensorRT-LLMKubernetesOpenTelemetry
Relevant experience matters more than knowing every tool listed.