Remote job
Lead Site Reliability Engineer
Job details
About this role
Role overview
Lead high-severity incident response for a large cloud software platform, coordinating technical, product, and security teams through complex service disruptions. This individual-contributor role combines incident command, stakeholder communication, operational analysis, and reliability improvement across a globally distributed environment.
Responsibilities
- Direct incident commanders and coordinate response to critical technology, product, and security incidents. - Communicate clearly with technical teams, business stakeholders, and executive leadership during challenging events. - Act as a subject-matter expert for incident-management practices and help strengthen the wider service-excellence program. - Apply engineering discipline, automation, and operational best practices to improve availability, reliability, and scalability. - Review incident data for anomalies, correlations, and recurring trends, then recommend improvements. - Participate in a continuous 24x7x365 rotational coverage schedule.
Requirements
- At least 12 years of experience in incident management or site reliability engineering. - Demonstrated leadership of major incidents and other high-severity operational situations. - Experience using incident-management tools and strong programming ability. - Advanced troubleshooting and problem-solving skills in a 24x7x365 environment. - Strong judgment, decision-making, and problem-identification capabilities. - Excellent written and verbal communication, including the ability to present effectively to business audiences.
Benefits and work setup
- Remote position, with most work performed from a designated remote location. - Individual-contributor role focused on operational leadership and reliability expertise.