Remote job
Member of Engineering (Inference Infrastructure)
Job details
About this role
Role overview This role sits on the compute team, partnering closely with the inference team to push the throughput and latency of large-model inference workloads. The work blends GPU workload scheduling, serving optimization, and the systems engineering needed to keep high-throughput data planes healthy under pressure. It is a deeply technical position aimed at compounding research and deployment velocity.
Responsibilities - Build and refine GPU workload scheduling and inference serving systems used for evals and reinforcement learning - Improve inference throughput and latency in collaboration with the inference team - Design and operate distributed systems, schedulers, control planes, or high-throughput data planes - Work hands-on with Kubernetes internals, including controllers, informers, and operators, rather than treating it as a deployment target - Champion observability and debuggability so production issues can be navigated quickly - Continuously raise the bar for systems performance and architectural velocity
Requirements - Strong programming skills in Go or a comparable systems language - Systems engineering background across distributed systems, schedulers, control planes, or high-throughput data planes - Production experience with Kubernetes internals, including controllers, informers, and operators - A bias toward observability and debuggability when designing production systems - Ability to thrive in a fast-paced, mission-driven environment with low-ego, collaborative teammates
Nice to have - Experience building systems that serve large-scale inference requests
Benefits and work setup - Fully remote role with flexible hours, based in the EMEA region - Roughly 37 days per year of vacation and holidays - Health insurance allowance for the employee and dependents - 16 weeks of flexible, full-pay parental leave - Well-being, always-be-learning, and home office allowances - Company-provided equipment - Regular in-person team gatherings, including monthly meetups in Paris and an annual longer off-site