Remote job
B2B Site Reliability Engineer
Job details
About this role
Role overview A remote Site Reliability Engineer role based in Poland, partnering with product engineering teams to balance delivery speed with the reliability of customer-facing services. The position sits at the intersection of engineering, product, customer success, and technical support, blending hands-on production operations with the design of automation and agentic tooling that improves reliability across the organization. Occasional in-person collaboration at a local office may be required for team events.
Responsibilities - Partner with engineering teams to define service-level objectives, error budgets, and supporting indicators, and use these measures to guide prioritization and reliability investments. - Investigate complex production issues end-to-end across application, data, infrastructure, and network layers, using AI to correlate logs, metrics, and code and to validate hypotheses. - Produce clear technical documentation, runbooks, architecture notes, postmortems, and proofs of concept aimed at both technical and non-technical audiences and structured for reuse by engineers and AI tools. - Identify systemic sources of toil and lead efforts to eliminate them through automation, AI agents, tooling, and process improvement. - Establish the conditions for AI agents to operate reliably, including repository context, well-specified tasks, integrations such as MCP servers, and the tests and guardrails needed to trust AI-authored changes. - Drive cross-team collaboration on reliability initiatives, mentor engineers on SRE practices, and advise senior leadership during critical customer escalations.
Requirements - At least 5 years of experience in software engineering, SRE, or production operations roles. - Strong production troubleshooting skills across the stack, using tools such as profilers, heap and thread dumps, query plans, traces, logs, and metrics to diagnose issues from first principles. - Hands-on experience operating production services on AWS (EC2, S3, EKS, RDS/Aurora, CloudFront). - Working knowledge of observability tooling such as Grafana, Prometheus, or LogicMonitor. - Practical experience writing infrastructure as code and automation in a general-purpose language such as Python, Go, or Java to a production standard. - Demonstrated judgement about applying AI across SRE work, including high-stakes areas like production access and sensitive data, and experience using agentic development tools such as Claude Code, Cursor, or GitHub Copilot. - Experience working within an Agile development framework and producing documentation for varied audiences.
Benefits and work setup - Fully remote role open to candidates based in and with the legal right to work in Poland. - Occasional in-person collaboration at a nearby office for events that benefit from face-to-face work. - Culture described as flexibility-oriented, trust-based, and supportive of work-life balance, with an emphasis on continuous learning and personal growth.