Remote job
Infrastructure / SRE Engineer
Job details
About this role
Role overview An infrastructure and site reliability engineering role responsible for keeping AI agent services scalable, performant, and highly available. The position spans application-level optimization, AWS architecture, CI/CD, observability, and incident response, with significant ownership on a small team.
Responsibilities - Build and maintain infrastructure on AWS ECS and related services. - Manage and optimize Supabase (PostgreSQL) for performance and reliability. - Design and implement CI/CD pipelines with GitHub Actions for fast, reliable deployments. - Set up monitoring, alerting, and observability across all services. - Optimize application and infrastructure performance for scale, targeting 99.99% uptime. - Review code for scalability and performance issues, automate provisioning, and lead incident response with post-mortems.
Requirements - Strong experience with AWS ECS and container orchestration. - Proficiency with Supabase or PostgreSQL performance tuning and optimization. - Experience building and maintaining CI/CD pipelines using GitHub Actions. - Solid understanding of containerization with Docker. - Familiarity with monitoring and observability tools such as Datadog, Prometheus, Grafana, or CloudWatch. - Comfort debugging production issues under pressure, with strong scripting skills in Python, Bash, or similar.
Nice to have - Experience scaling ECS services to handle high traffic. - Background in security best practices and compliance. - AWS cost optimization experience. - Deep knowledge of PostgreSQL and Redis optimization. - Experience with infrastructure-as-code tools such as Terraform or CloudFormation. - Open source contributions to infrastructure tooling.