Remote job
Production Support Engineer
Job details
About this role
Role overview
A small, high-trust reliability team is hiring a Production Support Engineer to safeguard a revenue-critical sales platform. The role sits between end users (partners and call centers) and engineering, triaging incidents and translating technical realities into clear updates. It is not a development position; the bar is calibrated for investigation and communication rather than implementing code fixes.
Responsibilities
- Monitor platform health and triage incoming alerts, distinguishing genuine outages and degradation from routine bugs and defects. - Investigate incidents through logging tools, reading API calls, responses, log data, and browser developer output to localize failures. - Own the incident lifecycle end-to-end, including stakeholder communication, post-mortems, and corrective-action follow-up. - Proactively notify stakeholders of critical issues, flagging SLA risk and maintaining status updates through email, phone, or ticketing. - Cross-reference tickets across multiple systems and shepherd defects through the full lifecycle to closure. - Translate technical findings into plain language for partners and call center audiences while managing expectations during resolution. - Maintain the incident queue in Jira and help prioritize bugs inside engineering sprint cycles. - Join on-call rotations after ramp, responding to alerts through OpsGenie inside defined SLA windows.
Requirements
- At least two years troubleshooting applications, servers, or infrastructure environments. - At least two years communicating clear status updates to stakeholders at varying levels of seniority. - Working knowledge of APIs, including interpreting calls and responses, with hands-on experience in Postman or similar tooling. - Practical familiarity with observability platforms such as Splunk, Datadog, or Sumo Logic. - SQL skills for ad hoc troubleshooting and reporting. - Comfort reading HTML and JSON and using browser developer tools for investigation. - Bachelor's degree in a related field or equivalent practical experience. - Availability for U.S. Eastern business hours (9 AM – 6 PM ET), with Eastern time strongly preferred for onboarding and on-call coordination.
Nice to have
- AWS knowledge at Cloud Practitioner level or above, plus familiarity with Git in team settings. - Exposure to infrastructure-as-code concepts such as Terraform. - Prior use of OpsGenie or comparable alerting platforms. - Experience applying AI tooling to investigation and troubleshooting workflows. - Background in on-call rotation structures, incident severity frameworks, and live partner or call center communications. - Familiarity with travel, hospitality, or high-volume transactional platforms.
Benefits and work setup
- A structured six-month onboarding ramp toward full self-sufficiency, with on-call rotations starting only after readiness and manager backup during early shifts. - An active alert window of 8 AM–1 AM Eastern, with overnight suppression periods built in. - SEV-1 incidents are rare, occurring roughly once per quarter or less, with most work happening during business hours. - An emphasis on longevity and growth within a tight-knit team that values genuine interest over short tenure.