Remote job
ML Engineer
Job details
About this role
Role overview This remote, full-time position focuses on training, fine-tuning, and serving compact specialized models that power a routing layer for AI inference. The work centers on making inference cheaper, faster, and higher quality than calling external model providers directly, by owning models end-to-end and squeezing the cost, latency, and quality trade-off curve to its frontier.
Responsibilities - Own the lifecycle of small specialized models covering use cases such as semantic search, intent classification, and content moderation. - Build training data pipelines that convert live production traffic into supervised signal while defaulting to privacy-respecting data practices. - Push the cost, latency, and quality trade-off curve from "good" to "obviously the best choice" through systematic optimization. - Evaluate and integrate techniques like distillation and quantization to reduce serving cost without sacrificing quality. - Read and triage recent research quickly, deciding which papers are worth implementing in production.
Requirements - Proven track record of shipping at least one model into production traffic and an understanding of what breaks at scale. - Strong intuition for distillation, quantization, and the parts of the inference stack that materially affect latency. - Ability to read a research paper and reach a go-or-no-go implementation decision in roughly 20 minutes. - Comfort working independently in a remote environment on full-time ML engineering problems.
Nice to have - Background in designing or operating high-throughput, low-latency inference systems. - Experience building privacy-preserving data pipelines from production traffic.
Benefits and work setup - Fully remote, full-time role with compensation in the $150,000–$400,000 USD range. - Structured interview process including a coding screen, two technical interviews, and a 30-minute founder conversation, with an offer typically extended within five business days of completing the loop.