Remote job
Senior Site Reliability Operations Engineer - Finance
Job details
About this role
Role overview A Site Reliability Operations Engineer is needed to safeguard 24/7 availability of mission-critical backend systems inside a regulated financial services environment. The position blends incident command, hands-on technical troubleshooting, infrastructure project leadership, and clear communication across engineering, vendors, and executives. The schedule is a swing shift running 2:00 PM to 10:30 PM PST.
Responsibilities - Monitor the health of a multi-platform IT estate with AWS CloudWatch, New Relic, Nagios, and SumoLogic, tuning alert thresholds to reduce noise and surface real issues early. - Act as an escalation point for complex problems, troubleshooting across Linux/UNIX, Windows, virtual servers, and virtual desktop infrastructure. - Coordinate, automate, and execute deployments through Jenkins, GitLab, or comparable CI/CD tooling, with a focus on zero-downtime releases. - Partner with application developers, third-party vendors, and internal incident management to resolve issues and drive continuous improvement. - Lead medium-to-large infrastructure initiatives such as cloud migrations, platform upgrades, and performance tuning. - Maintain standard operating procedures in the team knowledge base and oversee enterprise backup operations using CommVault, Veeam, and AWS Backup.
Requirements - 5+ years working in a Site Reliability Operations, NOC, or cloud infrastructure role with hands-on experience deploying full-stack applications. - Strong administration and troubleshooting skills across Windows and UNIX/Linux, including log analysis, performance diagnostics, and virtual server or desktop management. - Practical experience with AWS services (storage, compute, networking) and enterprise monitoring stacks such as CloudWatch, New Relic, Nagios, and SumoLogic. - Scripting or programming ability in PowerShell, Python, or bash for automation and operational optimization. - Familiarity with CI/CD platforms (Jenkins, GitLab), ITSM tools (ServiceNow, Jira), and enterprise backup solutions. - Clear verbal and written communication skills, with a track record of bridging technical teams, executives, and external vendors.
Nice to have - Advanced AWS certifications. - Exposure to AI/ML tooling for infrastructure monitoring or predictive analytics. - Experience in ITIL-aligned environments or formal change and incident management frameworks. - Bachelor's degree in Computer Science, Information Technology, or a related field.
Benefits and work setup - Fully remote position; only a laptop and reliable internet connection are required. - Compensation paid in USD at market-leading rates. - Paid time off policy supporting rest and recharging. - Results-oriented culture with flexible scheduling and high autonomy. - Opportunity to collaborate on high-impact initiatives for major U.S. companies within a large, globally distributed engineering community.