Remote job
Applied AI & ML Systems Engineer
Job details
About this role
Role overview
Build sub-100ms inference pipelines, custom agentic workflows, fine-tuning infrastructure, and high-performance vector retrieval architectures powering enterprise AI systems. The position bridges cutting-edge generative AI research and deterministic, low-latency production systems that must perform under heavy concurrent loads.
Responsibilities
- Engineer, optimize, and deploy high-throughput LLM inference architectures using vLLM, TensorRT-LLM, and Triton Inference Server. - Architect production-grade retrieval-augmented generation (RAG) pipelines and low-latency vector databases such as Qdrant or Milvus. - Implement custom model fine-tuning (LoRA or QLoRA), evaluation harnesses, and synthetic dataset pipelines. - Design agent orchestration runtimes with strict determinism, guardrails, and telemetry. - Profile and eliminate GPU memory bottlenecks while optimizing inference throughput.
Requirements
- 4+ years of engineering experience with a strong focus on applied machine learning and backend systems. - Strong proficiency in Python, PyTorch, and CUDA acceleration primitives. - Practical experience optimizing inference latencies through quantization, KV-cache management, or speculative decoding. - Solid foundation in distributed systems, asynchronous API design, and production monitoring. - Curiosity for translating cutting-edge generative AI research into production reality.
Nice to have
- Experience deploying models on Kubernetes with GPU operator acceleration. - Familiarity with Triton Inference Server C++ custom backends. - Published research or benchmark evaluations in LLM inference optimization.
Benefits and work setup
- Remote, full-time, globally distributed. - Compensation of $195,000 – $255,000 plus profit share. - Asynchronous working rituals with an engineering pod structure.