Remote job
Senior Incident & Problem Manager
Job details
About this role
Role overview A Senior Incident and Problem Manager is needed to take end-to-end ownership of incident management for an AI-driven personalization platform used by global brands. The role is hands-on and operational, combining live incident leadership with the design of a unified, product-wide process and the building of a problem management practice from scratch.
Responsibilities - Own incident management end to end and personally lead live incidents as Incident Manager, including war-room and bridge leadership through to restoration. - Design and implement a single, unified incident process across the product portfolio. - Own customer-facing incident communications, including status page updates and concise briefings for executives and internal stakeholders. - Run post-incident reviews, document learnings, assign corrective actions with owners and timelines, and drive them to closure. - Establish problem management from the ground up, including root cause analysis, a known-error database, and a path to prevent recurring incidents. - Partner with Engineering, Customer Success, and Product, and enable teams through training, documentation, and clear process guidance. - Build deep product proficiency at both a functional and technical level so incidents can be led with real product judgement. - Track and report on metrics such as MTTR, repeat-incident rate, and SLA compliance to drive continuous improvement. - Operate daily in Jira, Zendesk, and PagerDuty, and own the status page as the customer-facing source of truth. - Participate in a 24/7 on-call rotation.
Requirements - Five or more years of experience in incident management, major incident coordination, technical operations, or SaaS operations. - Willingness and availability to participate in a 24/7 on-call rotation as a core condition of the role. - Proven ability to lead war rooms and incident bridges under pressure with calm, structured, decisive execution. - Strong customer-facing communication, including translating technical disruption into clear status updates and stakeholder briefings. - Ability to deliver concise executive briefings on impact, risk, and recovery during critical incidents. - High ownership, hands-on approach, and a track record of building or maturing incident processes rather than only operating them. - Excellent influencing skills and the ability to align teams without formal authority. - Hands-on experience with ITSM and incident tooling such as Jira, Zendesk, and PagerDuty. - Strong analytical skills, including leading root cause analysis and distinguishing symptoms from causes.
Nice to have - Functional leadership of other incident and problem managers without direct people management. - Background in SaaS, AI, or developer-platform operations.
Benefits and work setup - Fully remote within the Czech Republic or Slovakia, with working hours centered on Central European Time. - Comprehensive benefits including generous paid time off, flexible work-from-home policy, parental leave, wellness programs, professional development budget, employee assistance, meditation and sleep app subscription, quarterly company-wide days off, employee resource groups, and equity or performance bonus participation.