Remote job
Technical Enterprise Incident Manager
Job details
About this role
Role overview This remote position centers on leading enterprise-scale incident response for a cloud-hosted production environment, coordinating across infrastructure, application, network, security, and vendor teams to restore services quickly. The role blends hands-on technical work with operational leadership, driving both immediate incident resolution and long-term reliability improvements. It is suited to an experienced operator who can bridge technical depth with executive-level communication.
Responsibilities - Lead and coordinate major incident bridge calls, driving rapid service restoration while maintaining accurate timelines and executive updates - Prioritize incidents based on business impact and operational risk, managing escalations and engaging leadership when needed - Facilitate post-incident reviews, produce high-quality root cause analysis documentation, and track corrective actions to completion - Drive automation of detection, triage, and response workflows to reduce manual effort and shorten resolution times - Partner with application teams to implement root-cause remediations and develop monitoring, alerting, and dashboarding capabilities - Maintain runbooks, service dependency maps, and incident management procedures; ensure accurate ticket documentation within the ITSM platform - Analyze operational trends and KPIs to identify reliability risks and support resiliency strategies such as redundancy, failover, and capacity planning
Requirements - U.S. citizenship with the ability to obtain and maintain a Public Trust clearance - Bachelor's degree with five years of relevant experience, or a high school diploma with nine years of relevant experience - Five or more years in cloud incident management, operations engineering, NOC, SRE, application, or production support roles - Demonstrated experience leading enterprise major incident response in a 24x7 operational environment - Working knowledge of ITIL incident and problem management processes - Three or more years of hands-on experience with AWS cloud services - Familiarity with observability platforms such as Datadog, Cloudcraft, or comparable tools, plus an enterprise ITSM platform
Nice to have - Background in Site Reliability Engineering or cloud platform DevOps environments - Experience supporting federal, healthcare, financial, or other highly regulated sectors - Hands-on familiarity with Windows and Linux servers, networking fundamentals, multi-cloud platforms (AWS, Azure, or GCP), load balancers, proxies, DNS, and firewalls
Benefits and work setup - Fully remote position scheduled 8:00 a.m. to 5:00 p.m. Eastern Standard Time, with participation in a 24x7 on-call rotation for incident management