Remote job
DevOps Engineer IV - US
Job details
About this role
Role overview Seeking a senior infrastructure engineer to own complex, cross-cutting cloud systems and CI/CD pipelines, including LLMOps workflows that support AI/ML model deployment. The position sets technical direction within the area, mentors engineers at multiple levels, and serves as a primary incident responder for high-impact issues. It is a hands-on leadership role for someone who combines deep architecture skills with mentorship and stakeholder communication.
Responsibilities - Architect and own end-to-end delivery of complex infrastructure and CI/CD systems, including LLMOps pipelines spanning multiple services. - Lead design reviews, set technical direction for the area, and mentor junior and mid-level engineers. - Drive incident response and root-cause analysis for high-impact infrastructure issues within the domain. - Help secure AI endpoints and agentic tools against prompt-injection and data-leakage risks for owned systems. - Contribute to and help enforce security and compliance controls across infrastructure. - Translate stakeholder needs into the infrastructure and AI-ops roadmap and occasionally present technical work to cross-functional groups of 5–10 people. - Participate in an on-call rotation that includes after-hours and weekend support.
Requirements - Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience. - 6–9 years of experience in technology, with significant infrastructure ownership. - 4–6 years of hands-on experience architecting and optimizing cloud infrastructure and CI/CD systems. - Advanced expertise in cloud infrastructure, containerization and orchestration, and infrastructure-as-code, with a track record of designing highly available, cost-optimized architectures. - Ability to independently design CI/CD systems and deployment pipelines, including LLMOps pipelines for AI/ML models. - Deep understanding of AI infrastructure operational signals such as latency, token cost, and drift, and the ability to design monitoring and alerting around them. - 1+ year informally mentoring engineers or leading design reviews.
Nice to have - Cloud Professional-level certification. - Deep experience with LLMOps at scale, including GPU provisioning and vector database hosting.
Benefits and work setup - Compensation varies by geographic market and is based on factors including location, job-related knowledge, skills, and experience. - The package may include incentive compensation such as annual bonus or incentives, equity awards, and an Employee Stock Purchase Plan (ESPP). - On-call rotation participation is required, including after-hours and weekend support.