Remote job
Senior Automation & Observability Engineer
Job details
About this role
Role overview The Senior Automation & Observability Engineer plays a critical role in keeping enterprise infrastructure, applications, IoT platforms, and telemetry ecosystems healthy, observable, and continuously optimized. This position blends hands-on engineering with operational ownership, partnering across Infrastructure, Cloud, Network, Application, and Service Delivery teams to elevate reliability and proactive incident detection. It is a senior, cross-functional role suited to someone who thrives on turning complex monitoring signals into clear, actionable insight.
Responsibilities - Design, build, and maintain enterprise monitoring and observability solutions across infrastructure, applications, middleware, and IoT services. - Develop dashboards, alerts, and visualizations using Grafana, and tune data collection with Telegraf, Prometheus, and other agents. - Configure, validate, and troubleshoot Enterprise Logging & Telemetry (ELT) integrations and FOAK service onboarding across hybrid environments. - Analyze metrics, logs, traces, and events to detect performance bottlenecks, support SLO/SLA targets, and lead Root Cause Analysis. - Administer InfluxDB time-series databases, including retention, performance tuning, and capacity planning for telemetry workloads. - Monitor and support Linux/Windows servers, VMware, Citrix VDI, DNS, proxy, middleware, integration, and enterprise application platforms.
Requirements - Significant experience operating enterprise observability platforms such as IBM Instana, Grafana, and SolarWinds. - Strong working knowledge of telemetry pipelines, log ingestion, event correlation, and data quality practices. - Hands-on experience with Linux and Windows server environments, plus familiarity with VMware and Citrix VDI monitoring. - Proficiency with time-series databases (InfluxDB), data retention strategies, and dashboard development. - Demonstrated ability to perform RCA, troubleshoot infrastructure anomalies, and drive continuous improvement. - Solid collaboration skills across infrastructure, cloud, network, and application teams in a service delivery context.
Nice to have - Background in Service Reliability Engineering (SRE), infrastructure automation, or platform engineering. - Experience supporting FOAK or First Office Application Kit services and enterprise integration platforms.
Benefits and work setup - Remote-first working with flexibility to work from home or from office locations. - Unlimited paid days off, flexible scheduling, and a sabbatical leave program. - Health, dental, vision, disability, life/AD&D, and flexible spending account options, plus family-forming and parental leave benefits. - 401(k) with company match, education/student loan assistance, and a wellness program. - Annual bonus plan tied to company and individual performance, plus an equity appreciation program.