Remote job
Systems and Automation Engineer
Job details
About this role
Role overview This position focuses on owning the automation backbone that powers both client engagements and internal systems. The work centers on replacing fragile, one-off scripts with durable, observable workflows that are designed to be retried, audited, and handed off cleanly. It is a hands-on systems and platform role for someone who treats automation as a product rather than a side effect.
Responsibilities - Design and operate scheduled and event-driven workflows across runtimes such as GitHub Actions, Cloud Run jobs, AWS Lambda, and durable workflow engines, with retries and idempotency as first-class concerns. - Build typed integration scripts in TypeScript or Python that connect commerce platforms, content management systems, data warehouses, and messaging tools, with explicit contracts at every boundary. - Add LLM-driven steps where they earn their place (classification, extraction, summarization, content generation), paired with schema validation and human review gates. - Instrument every job with structured logs, run history, and actionable alerts that name the failing step and likely fix, and build guardrails such as dry-run modes, staged writes, rate limits, and kill switches. - Manage secrets, credential rotation, and least-privilege service accounts for everything that runs unattended, and own the operational hygiene that keeps unattended work safe. - Consolidate scattered one-off scripts into shared, documented tooling, write runbooks so any engineer can rerun, backfill, or pause a pipeline, and audit what is running, retire what is unused, and keep automation costs honest.
Requirements - Four or more years in systems, platform, DevOps, or backend engineering with genuine production-automation ownership. - Strong scripting in TypeScript or Python, working knowledge of shell, and the judgment to recognize when a script should be promoted to a service. - Hands-on experience with GitHub Actions and at least one cloud job runtime such as Cloud Run, Lambda, ECS, or Kubernetes CronJobs. - Comfort with Postgres, message queues, and the failure modes of third-party APIs, including the discipline to fix logging after a 2 a.m. incident so the next on-call inherits a calmer system.
Nice to have - Experience using LLM APIs in batch or pipeline contexts, including cost control and output validation. - Familiarity with workflow engines such as Temporal, Inngest, Airflow, or Prefect. - Exposure to commerce or content operations, for example Shopify, Sanity, or another headless CMS.