Remote job
SysOps Engineer – Monitoring & Cloud Operations
Job details
About this role
Role overview Operate and monitor infrastructure that supports production services, with a focus on availability, incident response, and recovery readiness. The work includes observability, operating system administration, backups, disaster recovery, failover exercises, capacity planning, and operational documentation across cloud and other environments.
Responsibilities - Maintain monitoring, alerts, dashboards, and service health checks using observability tools; track CPU, memory, disk, and processes. - Respond to incidents, troubleshoot service problems, investigate root causes, and support uptime and service-level objectives. - Administer Linux and Windows systems, including patching and performance tuning. - Manage backups and validate restores; execute disaster recovery plans and coordinate outage simulations. - Test failover and failback for critical services, monitor replication and data consistency with data teams, and maintain recovery runbooks. - Plan capacity, optimize infrastructure performance, and keep operational logs and documentation current.
Requirements - Relevant experience in systems or cloud operations, infrastructure support, SRE, or a related field; a relevant degree or equivalent practical experience. - Hands-on Linux and Windows administration experience. - Experience with monitoring and observability platforms such as New Relic, Prometheus, Grafana, or Datadog, and with a cloud platform such as AWS, Azure, or Google Cloud. - Knowledge of incident and problem management, root cause analysis, backups, disaster recovery, business continuity, and failover. - Experience managing virtual machines, cloud instances, or physical servers; familiarity with services and web servers such as systemd, Nginx, or IIS. - Strong troubleshooting, analytical, communication, documentation, and cross-team collaboration skills.
Nice to have - Experience supporting highly available, mission-critical production systems.