Remote job
Senior Site Reliability Engineer (Fedramp)
Job details
About this role
Role overview A senior engineering role focused on operating and continuously validating regulated cloud estates against Key Security Indicator (KSI) standards. The position owns a single operational capability end-to-end, including its automation, runbooks, telemetry design, and recovery procedure, and serves as the senior escalation point for the most complex incidents. It is fundamentally an operations role centered on keeping authorized environments available and demonstrably compliant over time.
Responsibilities - Own one operational capability (such as monitoring and alerting, remediation, or backup and recovery), including the automation and procedures that every managed environment inherits. - Design telemetry, log pipelines, service-level objectives, and alert quality for regulated environments, ensuring escalations are actionable and routed correctly. - Build and maintain the continuous-monitoring evidence pipeline so the authorized state of the estate can be demonstrated in machine-readable form each day. - Engineer backup and recovery procedures with tested recovery objectives and the automation that executes them during outages. - Serve as the senior on-call escalation, diagnosing beyond the runbook, resolving incidents, and updating the runbook so the next responder does not repeat the diagnosis. - Lead incident and problem management for the capability, including incident command during major events, blameless post-incident review, and corrective action. - Reduce operational toil through infrastructure-as-code, pipelines, and scripting, choosing what to standardize across the estate.
Requirements - Significant hands-on experience operating production cloud or regulated environments at scale. - Strong infrastructure-as-code, CI/CD pipeline, and scripting skills. - Proven ability to design observability, alerting, and escalation systems that hold up under service-level commitments. - Track record of leading incident response and post-incident improvement in complex systems. - Experience working against contractual SLAs and regulatory controls. - Comfort owning a capability independently and mentoring surrounding engineers.
Nice to have - Recognized depth in a specialty such as SIEM and log pipelines, observability platforms, vulnerability and patch management at scale, or disaster-recovery engineering. - Background operating a monitored compliance environment against contractual service levels. - Certifications such as cloud security or DevOps specialty credentials, CISSP, or GIAC. - Experience supporting clients from within a professional services or managed services organization. - Prior involvement in renewals, expansions, and pre-sales technical solutioning for managed services. - Familiarity with configuration baseline standards such as CIS Benchmarks and DISA STIG. - Familiarity with frameworks such as FedRAMP, FISMA, HIPAA, HITRUST, or PCI.
Benefits and work setup - Reported salary range of approximately $85,000 to $141,000 per year, with eligibility to participate in annual incentive, commission, and recognition programs based on role, location, and experience. - Flexible work model that lets employees choose when and where they work most effectively, including remote and office-based options. - Paid parental leave and flexible time off. - Certification and training reimbursement. - Digital mental health and wellbeing support membership. - Comprehensive insurance options. - Employee resource groups and in-person and virtual events. - Employer describes itself as an equal opportunity employer committed to pay equity, reasonable accommodation, and non-discrimination on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.