Remote job
Staff Software Engineer, Platform Infrastructure
Job details
About this role
Role overview A platform infrastructure group is hiring a Staff Software Engineer to establish Site Reliability Engineering as a first-class discipline across the internal platform. The engineer will design reusable service templates, observability standards, and infrastructure patterns that dozens of engineering teams rely on to ship safely and quickly. This is a senior, hands-on technical leadership role combining platform engineering with SRE ownership over a Go-based stack on Google Cloud.
Responsibilities - Define and roll out SLO and SLI frameworks, incident response practices, and reliability standards across the platform group. - Set the observability baseline using Prometheus metrics, OpenTelemetry tracing, and structured logging so every team has actionable signals. - Build and maintain paved-path infrastructure: Go APIs, CLIs, Cloud Run services, GKE workloads, and Terraform modules with reliability built in from day one. - Lead migrations toward GCP-native managed services such as Secret Manager, Cloud SQL, and standardized CI pipelines on GitHub Actions and Cloud Build. - Apply security-first thinking to networking and identity (IAP, IAM, VPC design) and support compliance posture for SOC 2, ISO 27001, PCI DSS, and adjacent frameworks. - Mentor engineers, run full DevOps lifecycle work for the systems built, and participate in on-call after a typical three to six month ramp.
Requirements - 8+ years building and operating production systems with significant time in infrastructure, SRE, or platform engineering. - Demonstrated track record implementing SRE practices including SLO frameworks, incident management, on-call culture, and measurable reliability gains. - Strong hands-on proficiency in Go (1.21+) with Python as a secondary language. - Deep practical experience with GCP (Cloud Run, GKE, IAM, networking, managed services) or comparable clouds with willingness to ramp. - Hands-on Terraform experience building reusable modules, plus operational Kubernetes experience on GKE or equivalent. - Production experience with Grafana, Prometheus, and OpenTelemetry, paired with practical IAP, IAM, and VPC work inside compliance frameworks.
Nice to have - Platform engineering philosophy: thinking in patterns and self-service so teams get great templates rather than one-off solutions. - Technical leadership experience setting direction for a platform, translating ambiguous reliability goals into architecture, and influencing peers outside direct reporting lines. - Clear written and verbal communication for articulating reliability risk, architectural decisions, and postmortems to both engineers and non-technical stakeholders. - A team-oriented mindset where success is measured by collective outcomes rather than individual output.
Benefits and work setup - Remote within Ireland, collaborating primarily with teams in North America and Europe; some scheduling flexibility is expected for cross-timezone standups and incident response. - Compensation benchmarked to the Irish market, with compliance with all applicable Irish employment law including statutory leave entitlements. - Industry-competitive pay plus equity. - 28 days of holiday, private medical and dental coverage, life and critical illness insurance, and a workplace pension scheme. - Top-of-line equipment, an Employee Resource Platform, a monthly allowance for wellness and reading, and access to LinkedIn Learning. - Visa sponsorship is not available for this role.