Remote job
Support Engineer
Job details
About this role
Role overview Keep a large-scale Big Data platform and its data products running smoothly every day. The role spans the full incident lifecycle, from first alert and triage through root-cause analysis to permanent remediation and prevention, working across data pipelines and a mix of cloud and on-premises environments. It suits someone early in their infrastructure career who wants to grow toward a DevOps track while supporting both engineering and business-analysis teams.
Responsibilities - Own incident management end to end: detect, triage, diagnose, fix, and prevent recurrence. - Maintain environments and applications across AWS, on-prem servers, and Kubernetes clusters. - Provide day-to-day technical support to development and business-analysis teams. - Configure, tune, and maintain monitoring and alerting systems. - Use scripting and Linux tooling to automate routine operational tasks. - Contribute to the reliability of data pipelines feeding downstream data products.
Requirements - At least six months of experience in a Support Engineer or similar operational role. - Basic Linux administration skills and comfort working on the command line. - Hands-on experience with Git for version control. - Working knowledge of shell scripting (sh/bash). - Familiarity with network and diagnostic utilities such as curl, telnet, and dig. - Exposure to monitoring and alerting platforms like Prometheus, Grafana, or equivalents.
Nice to have - Configuration management experience with Ansible or Terraform. - Scripting in Python (preferred) or another high-level language. - Container experience with Docker and orchestration with Kubernetes. - Familiarity with CI/CD tools such as GitLab CI. - Exposure to the ELK stack for logging and search. - Understanding of Infrastructure-as-Code and GitOps principles. - Cloud platform experience, preferably AWS. - Working knowledge of relational databases such as PostgreSQL. - A clear growth path toward a DevOps engineering role.