Remote job
Senior Software Engineer, Site Reliability & Security
Job details
About this role
Role overview
This is a senior engineering position focused on operating a production platform that powers AI-driven voice technology used by enterprise and professional teams. The role blends site reliability engineering with platform security, owning availability, performance, and deployment safety for core services and data pipelines at meaningful scale.
Responsibilities
- Build and operate the monitoring, tracing, and alerting infrastructure that keeps production systems observable. - Drive platform security initiatives aimed at preventing incidents and hardening systems against threats. - Lead incident response and recovery efforts, including performing root cause analysis after active issues. - Run services with high availability and redundancy as primary design constraints. - Participate in an on-call rotation, responding to alerts and investigating issues across the stack. - Improve deployment pipelines so code changes are quick, simple, and safe to ship. - Identify novel ways to handle load and scale resource-intensive applications, especially inference and data workloads. - Partner with product engineers to help them write reliable code without slowing delivery.
Requirements
- 5+ years of professional experience with a modern programming language such as Go, TypeScript, or Python. - Strong grasp of computer science fundamentals including algorithms, data structures, and systems design. - Hands-on experience with cloud environments and container technologies, ideally AWS or GCP, Kubernetes, and Docker. - Practical knowledge of Infrastructure as Code using Terraform, OpenTofu, or Pulumi. - Familiarity with GitOps or continuous delivery tooling such as ArgoCD, Spacelift, or Terraform Cloud. - Ability to debug and resolve issues in complex production environments, including profiling applications and database performance. - Fluency working in a UNIX shell to analyze logs and handle common operational tasks. - Experience with monitoring tools like Grafana and Prometheus. - Must be a U.S. citizen or permanent resident and able to pass a pre-employment background check.
Nice to have
- Experience building and troubleshooting Kubernetes environments under production load. - Comfort profiling databases and tuning performance for data-intensive workloads.
Benefits and work setup
- Fully distributed team across the U.S. and flexible scheduling, with collaboration over chat and video. - Competitive compensation package including salary and equity. - Full medical, dental, and vision coverage. - 401(k) plan with company matching. - Generous paid time off policy and parental leave. - Annual learning and development stipend. - Home office stipend to support a remote workstation.