Remote job
Incident Manager - Senior [Customer IT Support]
Job details
About this role
Role overview A senior Incident Manager is needed to lead critical response efforts on a Customer IT Support team that coordinates high-impact outages across a fast-moving, customer-facing technology platform. The role blends hands-on incident command with process improvement, owning major incidents from detection through retrospective while shaping how the broader organization handles reliability and communication.
Responsibilities - Own end-to-end incident response, including detection, live coordination, resolution, and post-incident review, making time-critical decisions to limit customer impact - Take personal accountability for major incidents and drive them to closure with rigor and urgency - Coordinate cross-functional technical and business stakeholders globally, assigning roles and delegating tasks during active incidents - Lead customer-facing and internal communications throughout the incident lifecycle, including recovery messaging - Run blameless post-incident reviews, identify root causes, and convert findings into concrete process, monitoring, and alerting improvements - Mentor and train incident participants and duty team members, raising the overall maturity of incident handling across the company
Requirements - At least 3 years of experience leading significant incidents in a 24/7, high-load production environment - At least 5 years in engineering, product, or service delivery roles with increasing responsibility - Clear, structured communication skills, including the ability to brief executive leadership under pressure - Proven ability to coordinate multiple teams simultaneously and make decisive calls in ambiguous situations - Technical depth across public cloud, microservices, and modern monitoring/observability tooling - Comfort understanding a complex technology stack, sizing up fast-moving situations, and adjusting tactics on the fly
Nice to have - Background in fintech or in SRE-adjacent technical roles - Hands-on experience with AWS, Kubernetes, PostgreSQL, and Kafka - Familiarity with incident management platforms such as incidents.io, Rootly, or FireHydrant - Experience partnering with Legal, Communications, Privacy, and other non-technical functions during incidents
Benefits and work setup - Hybrid or fully remote within time zones from GMT-8 to GMT-3, with the option to work from an office in Mexico - Relocation support to Mexico covering the employee and their family - Healthcare coverage for employees - Education budget for language lessons, professional training, and certifications - Wellness budget for mental health and fitness activities - 20 days of annual leave plus paid sick leave - A culture that emphasizes innovative thinking, candid feedback, peer support, and shared recognition of wins