← Back to jobs

Remote job

Site Reliability Engineer ll

DevOps Full-time Permanent US (except California)

Job details

$100,000 Salary
US (except California) Eligibility
Lead Experience
Full-time Employment

About this role

Role overview A remote-first Site Reliability Engineer role supporting production healthcare systems, with occasional travel for onboarding and team events. The position bridges AWS cloud infrastructure, MERN stack applications, and large-scale data workflows, splitting effort roughly 60% on live operations and 40% on engineering automation to eliminate toil.

Responsibilities - Maintain continuous uptime, scalability, and security of AWS-hosted MERN applications and backend data architectures - Manage, optimize, and troubleshoot event-driven serverless architectures on AWS Lambda, focusing on cold-start mitigation, memory allocation, and execution timeouts - Monitor scheduled PySpark data workflows, execute standard operating procedures for large-scale ingestion, and rapidly triage, rerun, or patch failed jobs - Participate in a collaborative on-call rotation to triage, debug, and mitigate live application outages and data flow bottlenecks - Engineer automated workflows to eliminate repetitive tasks like manual data seeding, infrastructure provisioning, and routine PySpark recovery steps - Build specialized dashboards and alerts for Node.js event loops, PySpark execution stages, memory leaks, and pipeline anomalies, and lead blameless post-mortems

Requirements - 3+ years of hands-on experience operating multi-tenant, cloud-hosted, or cloud-native SaaS platforms at scale - Deep expertise operating AWS core services including Lambda, ECS/EKS, EMR or Glue, EC2, VPC networking, IAM, and CloudWatch - Professional competency in Python (including PySpark) and Node.js for automation scripts and data tooling - Experience managing distributed data orchestration pipelines, ETL tools, and message queues such as SQS/SNS or RabbitMQ - Strong understanding of the operational lifecycle of JavaScript/TypeScript applications, including memory management, asynchronous runtimes, and Node.js clustering - Practical experience managing, sharding, indexing, and optimizing production MySQL and Athena databases, plus infrastructure-as-code proficiency with Terraform or OpenTofu

Nice to have - 1+ year working within HIPAA-regulated environments and securing patient data at rest and in transit - 4+ years of software/systems experience with at least 1-2 years focused on live cloud operations and distributed data workflows

Benefits and work setup - Fully remote with approximately 5% travel - Medical, dental, vision, life, and disability insurance plus an Employee Assistance Program - 401K retirement plan and bonus eligibility

Skills detected in the listing

TypeScriptJavaScriptNode.jsPythonData EngineeringStakeholder ManagementAWSTerraform
Detected Sep 29, 2026
Last verified Sep 29, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight