Remote job
Systems Engineer
Job details
About this role
Role overview Join the platform support team of an AI and ML infrastructure organization as an early-career systems engineer with real ownership of production environments. The role splits between Internal Platforms, running the systems the company itself depends on, and Project Support, helping senior engineers on customer and government deployments. It is a fully remote, U.S.-based position that uses the same open-source building blocks internally and on customer projects as a path to grow into broader platform ownership.
Responsibilities - Own the internal request queue: triage incoming work, resolve what you can, escalate when needed, and ensure nothing is silently dropped. - Administer Google Workspace and identity systems, including accounts, groups, SSO integrations, security policy, and onboarding and offboarding runbooks. - Support provisioning and maintenance of cloud infrastructure, covering accounts, IAM roles, networking basics, resource tagging, and cost visibility. - Deploy and operate internal applications on Kubernetes using Helm and ArgoCD, with growing autonomy over time. - Maintain the company website and content publishing workflows, including build pipelines, deploys, and DNS. - Build and tune monitoring and alerting for owned systems, reducing noise instead of adding to it. - Handle customer-facing issues by reproducing problems, gathering context, executing runbooks, and tracking issues to closure. - Assist with hardening and compliance work, applying hardened configurations and gathering evidence for control implementation.
Requirements - Working knowledge of SQL and relational databases to pull data and build dashboards independently. - Hands-on Kubernetes experience using kubectl, Helm, and log and event inspection to debug pod issues. - Exposure to GitOps tooling such as ArgoCD or Flux, or CI/CD pipelines like GitHub Actions. - Some infrastructure-as-code familiarity with OpenTofu, Terraform, Pulumi, or Ansible. - Familiarity with identity and SSO standards such as Keycloak, OIDC, or SAML. - Experience with monitoring and observability tools like Prometheus, Grafana, or OpenTelemetry. - Comfort with containers and at least one major cloud provider such as AWS, Azure, or GCP. - Scripting ability in Python or Bash sufficient to automate repetitive work. - Strong written communication, since most of the role is conducted through tickets, pull requests, and documentation. - U.S. citizenship and eligibility to obtain and maintain a Secret security clearance.
Nice to have - Security+ or comparable certification, or coursework in security fundamentals. - Exposure to regulated or government environments, including RMF, NIST 800-53, FedRAMP, or restricted-network deployment. - Frontend or static site experience, particularly with Astro or a headless CMS. - Familiarity with the Python data science ecosystem or JupyterHub. - Open-source contributions of any size.
Benefits and work setup - Salary range of $85,000 to $120,000 USD, depending on experience and location. - Fully remote within the United States. - 100% employer-paid medical premiums for employees. - Self-managed paid time off with a minimum time-off requirement. - Asynchronous, globally distributed team.