Remote job
Software Engineer II, MLOps Framework
Job details
About this role
Role overview
This role sits on an MLOps framework and model conversion team that takes machine learning models from research into production on edge hardware for an autonomous driving program. The engineer owns the pipelines responsible for model conversion, compilation, benchmarking, and release across heterogeneous platforms, working closely with perception and safety counterparts to ensure every shipped model meets strict latency and accuracy requirements.
Responsibilities
- Design and implement model conversion and compilation pipelines using tools such as ONNX, TensorRT, and torch.compile for deployment on edge devices including NVIDIA Orin. - Maintain and evolve the model release registry, preserving traceability and reproducibility across model versions and target platforms. - Run rigorous latency benchmarking and model quality parity evaluations to validate safety-critical performance before release. - Compare metrics across hardware platforms to verify accuracy and latency compliance against release gates. - Communicate findings to model development teams and broader stakeholders, translating results into reliable, actionable follow-ups across the organization.
Requirements
- Bachelor's degree in Computer Science, Robotics, Electrical Engineering, or a related technical field plus around four or more years of relevant experience, or a Master's degree in a similar field plus roughly zero to three years of experience. - Extensive hands-on experience building model conversion and compilation pipelines with ONNX, TensorRT, and torch.compile, and performing latency benchmarking and quality parity validation. - Practical experience deploying and testing models on edge hardware such as NVIDIA Orin or comparable embedded platforms. - Experience maintaining a model release registry in a production environment. - Ability to interpret performance metrics across hardware platforms and validate models against strict latency and accuracy requirements.
Nice to have
- Expertise in model quantization techniques such as PTQ and QAT, and mixed-precision inference across INT8, FP8, FP4, BF16, and FP16. - Experience releasing multi-target models across heterogeneous platforms. - Familiarity with state-of-the-art autonomous driving perception algorithms, including temporal 3D object detection, BEV, and 3D occupancy networks, alongside multi-modal sensor fusion across vision, LiDAR, and radar. - C++ and/or CUDA kernel development experience.
Benefits and work setup
- Compensation package includes a bonus component and equity participation. - Medical, dental, and vision premiums covered for full-time employees. - 401(k) plan with an employer match. - Flexible scheduling and generous paid vacation available from the start date. - Life and accidental death and dismemberment insurance. - US pay range disclosed in the original listing.