Remote job
Site Reliability Engineer (Contract)
Job details
About this role
Role overview A vertically integrated logistics and technology company is hiring a contract Site Reliability Engineer to safeguard uptime for an AI-powered gate automation and computer vision platform running in real truck terminals across the United States. With platform load projected to grow roughly 10x over the next 18 months, reliability is now core to whether major logistics customers trust the system to keep their yards moving. You'll own incident response, monitoring, and proactive resilience across backend services, worker processes, and ML pipelines.
Responsibilities - Own reliability targets across backend APIs, worker services, applications, and computer vision pipelines, tracking MTD, MTM, and MTR and following through on root causes - Build out monitoring, alerting, and auto-remediation so on-call load scales with automation rather than headcount - Partner with internal teams developing agentic tooling to build agents that triage alerts and handle routine remediation - Harden and optimize GCP infrastructure (Cloud Run, Cloud SQL, GCS) for cost and performance under 10x load growth - Own database scale and performance: connection pooling, query optimization, indexing, read replicas, and capacity planning - Run blameless postmortems and drive fixes to root causes, not symptoms - Participate in an on-call rotation
Requirements - 4+ years in SRE, infrastructure, or backend engineering with production on-call ownership - Deep experience with a major cloud provider (GCP preferred) across compute, managed databases, object storage, and networking - Hands-on experience building observability stacks with tools like Grafana, Prometheus, Zabbix, or Datadog - Strong scripting and automation skills in Python, Bash, or similar - Comfort with containerized workloads (Docker) and CI/CD pipelines - Demonstrated track record of reducing incident volume or improving reliability metrics - Strong English communication skills, comfortable working async across technical and non-technical stakeholders
Nice to have - Experience with ML/data infrastructure, including training pipelines, model monitoring, or feature stores - Background building or integrating AI agents for operational automation such as alert triage and auto-remediation - Infrastructure-as-code experience with Terraform or similar tools - PostgreSQL performance tuning at scale - Exposure to physical/IoT systems including edge devices, cameras, and on-site hardware - Familiarity with bare-metal infrastructure in colocation environments, hardware monitoring, and high-availability design
Benefits and work setup - Full-time remote contract role, Latin America preferred - Embedded in core engineering sprints, standups, and Slack channels; employment is managed through an agency partner - Stack includes TypeScript/Node.js, Next.js, React, Apollo Server, Express, PostgreSQL, GCP, Docker, and Python for ML/data workloads