Remote job
Member of Technical Staff | Inference Platform
Job details
About this role
Role overview Own the systems that run machine learning models for both real-time API requests and large batch workloads. You’ll develop and operate inference infrastructure across cloud and customer-managed Kubernetes environments, with a focus on reliable releases, predictable performance, and efficient use of compute.
Responsibilities - Develop an online and batch inference runtime on Kubernetes, including its controllers and custom resources. - Run large inference batches as short-lived jobs, managing admission according to CPU, memory, and GPU capacity. - Improve model inference and feature-processing performance, and serve graph data efficiently. - Operate training, post-training, and fine-tuning jobs across cloud and customer-hosted environments. - Build autoscaling, GPU-serving, and performance improvements, using telemetry to track model behavior and resource use. - Improve availability, latency, batch throughput, cost per prediction, GPU utilization, and job completion reliability.
Requirements - Experience operating model serving or large-scale batch compute on Kubernetes. - Experience building Kubernetes controllers or operators. - Ability to profile and optimize data-intensive Python pipelines. - A practical focus on compute efficiency and operating costs. - Ability to write and review production-quality code and take responsibility for running it.
Nice to have - Production experience with Ray, Ray Serve, or KubeRay; admission control; or GPU serving and optimization. - Familiarity with Arrow, Parquet, Lance, or other columnar data formats. - Experience deploying to customer-managed Kubernetes, using GCP or AWS, or working in regulated environments.
Benefits and work setup Full-time, remote role listed for São Paulo. The work spans cloud and customer-hosted deployments.