Remote job
Staff Systems Software Engineer
Job details
About this role
Role overview
A staff-level systems software engineer is needed to build the tools, infrastructure, and quality systems that determine how fast and how reliably a studio ships gameplay and backend services. Rather than shipping player-facing features directly, this role builds the test harnesses, load and chaos tooling, device labs, and build and GPU clusters that prove the work of other teams is correct, fast, and ready. The position is remote and open to candidates in the US and Canada, partnering with cloud, gameplay backend, and cross-platform backend groups.
Responsibilities
- Build and own the studio's functional test automation framework, including runner, fixtures, game-state harness, reporting, and coverage instrumentation. - Design load, stress, and chaos testing systems that model realistic player behavior, simulate faults and partitions, and run safely against pre-production infrastructure. - Build app quality automation that captures frame time, memory, thermal, battery, and cold-start metrics automatically and gates them in CI. - Operate on-prem build clusters for distributed compilation, caching, and artifacts, plus on-prem GPU clusters with scheduling, multi-tenancy, and utilization tracking. - Define and publish developer velocity and quality metrics, and drive adoption across teams that do not report into the role. - Mentor engineers in testing strategy, performance analysis, and systems thinking, and raise the bar through code review and documentation.
Requirements
- Strong systems-level programming in Python, Golang, C or C++, Jenkins Groovy, and Linux shell scripting, plus fluent English communication. - Hands-on experience designing test frameworks, load and stress tooling (such as k6, Locust, Gatling, or JMeter), fault injection and chaos engineering, coverage instrumentation, and flake detection. - Proficiency in CPU, GPU, memory, and I/O profiling, statistical analysis of noisy benchmarks, regression detection, commit attribution, and observability. - Experience operating CI at scale with Jenkins or GitLab, Make or CMake, GCC or Clang, distributed build systems and caching, and large-binary workflows in Git, GitLab, or Perforce. - Comfort with Linux and macOS administration, bare-metal and on-prem cluster operations, Docker or Podman, Kubernetes, Terraform or OpenTofu, and hardware capacity planning. - Strong grasp of distributed systems failure modes, TCP/IP and UDP, resilience patterns, and disaster recovery planning.
Nice to have
- Mobile device farm operation, Android or iOS build and profiling toolchains, and GPU cluster scheduling using Slurm, Kubernetes device plugins, or MIG. - Game engine internals or engine-level test automation, deterministic simulation and replay testing, and live-service game experience or SaaS/PaaS equivalent. - Windows administration, Ansible, and additional languages such as TypeScript, Java, Kotlin, or Swift. - Experience designing AI-assisted developer tools, including agents, skills, or MCP servers.
Benefits and work setup
- Remote position open to candidates in the US and Canada. - Salary range of $162,500 to $219,000 USD annually, with the opportunity to earn an annual discretionary bonus. - Medical, dental, and vision coverage for employees and dependents starting on day one. - Paid time off, holidays, and a two-week winter break. - Pet insurance, compassionate leave, pre-tax wellness stipend, pre-tax work-from-home stipend, 401(k) with company match, mental health resources including Headspace and an EAP, discount portal, and support for professional development.