Remote job
Machine Learning Engineer
Job details
About this role
Role overview This full-time position focuses on the machine learning layer inside custom AI systems built for business operations, including document extraction, retrieval, classification, assistants with tool use, and internal tools. The team works remotely across the US and Australia and ships individual engagements on roughly six-week cycles. The role owns the ML end-to-end: problem framing, training or integration, offline evaluation, production integration, and post-launch quality. Foundation model training is explicitly out of scope.
Responsibilities - Collaborate with clients and engineering teammates to define ML problem statements, including inputs, outputs, latency/cost/accuracy constraints, and success criteria. - Design and ship systems for retrieval, extraction, classification, ranking, and tool-calling agents, choosing classical ML approaches where they outperform an LLM. - Build offline evaluation sets and harnesses, measuring quality, latency, and cost before any release reaches production. - Train, fine-tune, or distill models when a hosted API is the wrong fit; otherwise integrate existing models. - Embed models into Python-based production services and work across the stack in JavaScript or TypeScript. - Take ownership of production quality after launch through monitoring, failure analysis, regression tests, and follow-up iterations.
Requirements - Three or more years of experience shipping ML systems in production. - Strong Python skills with the ability to read and write JavaScript or TypeScript. - Hands-on experience with PyTorch, scikit-learn, or an equivalent training and evaluation stack. - Demonstrated ability to evaluate model quality across accuracy, precision/recall, hallucination, latency, and cost, and to iterate systems based on those results. - Working hours that overlap with Sydney, Austin, or Los Angeles.
Nice to have - Production LLM systems experience such as retrieval-augmented generation, tool calling, structured extraction, and evaluation harnesses. - Fine-tuning or model distillation work. - Monitoring and incident response for model-backed products. - Comfort turning an operations problem shared by non-engineers into a specification and dataset.