Remote job
Site Reliability Engineer ll
Job details
About this role
Role overview A remote-first Site Reliability Engineer role supporting production healthcare systems, with occasional travel for onboarding and team events. The position bridges AWS cloud infrastructure, MERN stack applications, and large-scale data workflows, splitting effort roughly 60% on live operations and 40% on engineering automation to eliminate toil.
Responsibilities - Maintain continuous uptime, scalability, and security of AWS-hosted MERN applications and backend data architectures - Manage, optimize, and troubleshoot event-driven serverless architectures on AWS Lambda, focusing on cold-start mitigation, memory allocation, and execution timeouts - Monitor scheduled PySpark data workflows, execute standard operating procedures for large-scale ingestion, and rapidly triage, rerun, or patch failed jobs - Participate in a collaborative on-call rotation to triage, debug, and mitigate live application outages and data flow bottlenecks - Engineer automated workflows to eliminate repetitive tasks like manual data seeding, infrastructure provisioning, and routine PySpark recovery steps - Build specialized dashboards and alerts for Node.js event loops, PySpark execution stages, memory leaks, and pipeline anomalies, and lead blameless post-mortems
Requirements - 3+ years of hands-on experience operating multi-tenant, cloud-hosted, or cloud-native SaaS platforms at scale - Deep expertise operating AWS core services including Lambda, ECS/EKS, EMR or Glue, EC2, VPC networking, IAM, and CloudWatch - Professional competency in Python (including PySpark) and Node.js for automation scripts and data tooling - Experience managing distributed data orchestration pipelines, ETL tools, and message queues such as SQS/SNS or RabbitMQ - Strong understanding of the operational lifecycle of JavaScript/TypeScript applications, including memory management, asynchronous runtimes, and Node.js clustering - Practical experience managing, sharding, indexing, and optimizing production MySQL and Athena databases, plus infrastructure-as-code proficiency with Terraform or OpenTofu
Nice to have - 1+ year working within HIPAA-regulated environments and securing patient data at rest and in transit - 4+ years of software/systems experience with at least 1-2 years focused on live cloud operations and distributed data workflows
Benefits and work setup - Fully remote with approximately 5% travel - Medical, dental, vision, life, and disability insurance plus an Employee Assistance Program - 401K retirement plan and bonus eligibility