← Back to jobs

Remote job

Senior Automation & Observability Engineer

DevOps Full-time Permanent United States

Job details

Not specified Salary
United States Eligibility
Senior Experience
Full-time Employment

About this role

Role overview The Senior Automation & Observability Engineer plays a critical role in keeping enterprise infrastructure, applications, IoT platforms, and telemetry ecosystems healthy, observable, and continuously optimized. This position blends hands-on engineering with operational ownership, partnering across Infrastructure, Cloud, Network, Application, and Service Delivery teams to elevate reliability and proactive incident detection. It is a senior, cross-functional role suited to someone who thrives on turning complex monitoring signals into clear, actionable insight.

Responsibilities - Design, build, and maintain enterprise monitoring and observability solutions across infrastructure, applications, middleware, and IoT services. - Develop dashboards, alerts, and visualizations using Grafana, and tune data collection with Telegraf, Prometheus, and other agents. - Configure, validate, and troubleshoot Enterprise Logging & Telemetry (ELT) integrations and FOAK service onboarding across hybrid environments. - Analyze metrics, logs, traces, and events to detect performance bottlenecks, support SLO/SLA targets, and lead Root Cause Analysis. - Administer InfluxDB time-series databases, including retention, performance tuning, and capacity planning for telemetry workloads. - Monitor and support Linux/Windows servers, VMware, Citrix VDI, DNS, proxy, middleware, integration, and enterprise application platforms.

Requirements - Significant experience operating enterprise observability platforms such as IBM Instana, Grafana, and SolarWinds. - Strong working knowledge of telemetry pipelines, log ingestion, event correlation, and data quality practices. - Hands-on experience with Linux and Windows server environments, plus familiarity with VMware and Citrix VDI monitoring. - Proficiency with time-series databases (InfluxDB), data retention strategies, and dashboard development. - Demonstrated ability to perform RCA, troubleshoot infrastructure anomalies, and drive continuous improvement. - Solid collaboration skills across infrastructure, cloud, network, and application teams in a service delivery context.

Nice to have - Background in Service Reliability Engineering (SRE), infrastructure automation, or platform engineering. - Experience supporting FOAK or First Office Application Kit services and enterprise integration platforms.

Benefits and work setup - Remote-first working with flexibility to work from home or from office locations. - Unlimited paid days off, flexible scheduling, and a sabbatical leave program. - Health, dental, vision, disability, life/AD&D, and flexible spending account options, plus family-forming and parental leave benefits. - 401(k) with company match, education/student loan assistance, and a wellness program. - Annual bonus plan tied to company and individual performance, plus an equity appreciation program.

Skills detected in the listing

PythonAWSGCPAzureDockerKubernetes
Detected Sep 18, 2026
Last verified Sep 18, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight