Remote job
Senior Site Reliability Engineer
Job details
About this role
Role overview A senior infrastructure role on a team building and operating the next generation of a satellite imaging platform that goes beyond traditional cloud delivery to support on-premises customer deployments. The work centers on architecting reproducible, air-gapped-capable systems for end-to-end imaging operations across cloud and customer environments.
Responsibilities - Build and deploy computing services and infrastructure for a satellite operations and image processing platform running in customer environments. - Architect novel systems for air-gapped deployments at scale within a high-impact team. - Clarify and surface requirements from ambiguous use cases defined by cross-functional stakeholders, including internal users and external customers. - Own deployment, service orchestration, and documentation operations for cross-platform stakeholders. - Scale architecture while preserving availability, improving reliability and scalability by studying failure modes and writing tests. - Participate in on-call rotations to ensure operational excellence.
Requirements - 6+ years building services that leverage cloud-native infrastructure and tooling, plus a Bachelor's degree in Computer Science or a similar field. - Experience deploying and maintaining bare-metal and cloud Kubernetes using tools such as Talos, RKE2, Proxmox, or k3s. - Proficiency with Terraform, Ansible, Helm, Kustomize, or similar infrastructure-as-code and GitOps tooling. - Experience with CI/CD tooling such as Jenkins, GitLab CI/CD, Argo CD, or CircleCI. - A track record of building, releasing, and supporting highly available, consistently performant services, plus knowledge of hardware and network implications of on-prem compute. - Experience with platform optimization including resource management and cluster tuning in constrained environments, distributed systems observability using Alloy, Prometheus, Grafana, or OpenTelemetry, and advanced skills in Python, Bash, and related tooling.
Nice to have - Experience with CUDA-based GPU programs. - Security expertise in sensitive environments.
Benefits and work setup Full-time remote position based in the United States or Canada, with employees near an office expected to work from that office three days per week. Global team with offices across North America and Europe and a culture that emphasizes iteration and putting team members first.