Remote job
Private Cloud & Infrastructure Operations Engineer
Job details
About this role
Role overview A Private Cloud and Infrastructure Operations Engineer is being hired to own day-to-day operations of a private cloud platform and contribute to broader compute and storage operations across data center environments. The role is hands-on and consumer-facing, sitting between the platform and its users to keep services reliable, efficient, and well governed.
Responsibilities - Monitor platform health, utilization, and capacity across all private cloud environments. - Manage compute provisioning, configuration, and lifecycle operations. - Track resource utilization, deliver reporting, and drive efficiency and accountability with consumers. - Manage quotas, resolve capacity issues, and plan for growth. - Provide first-line operational support, triage incidents, apply known fixes, and escalate appropriately. - Enforce operational standards, access controls, and governance policies. - Manage environment intake and provisioning for consumers. - Support compute and storage infrastructure operations, including server lifecycle, distributed storage, and shared services. - Maintain operational documentation and runbooks. - Participate in an on-call rotation for infrastructure services. - Contribute to infrastructure efficiency and automation initiatives.
Requirements - Three or more years operating cloud or virtualized infrastructure at scale. - Strong Linux systems administration experience with RHEL, CentOS, or Ubuntu. - Hands-on experience with distributed storage such as Ceph or with SAN operations. - Experience with monitoring stacks, dashboards, alerting, and capacity planning. - Virtualization experience with OpenStack, KVM, VMware, or similar platforms. - Kubernetes cluster operations, including node management and cluster lifecycle. - A governance mindset, comfortable enforcing standards, reporting compliance, and holding consumers accountable for waste. - Ability to author and maintain clear runbooks, procedures, and operational documentation. - Strong communication skills, with the ability to translate technical state into clear stakeholder reporting.
Nice to have - Hands-on Taikun operational experience. - Server hardware lifecycle experience, including racking, commissioning, decommissioning, and IPMI, iDRAC, or iLO. - Automation experience with Ansible, Terraform, or scripting languages. - ITSM or ITIL awareness across incident, change, and problem management. - Team leadership or mentoring experience. - FinOps or cost attribution experience.
Benefits and work setup - Generous paid time off policy and unplugged days. - Flexible work-from-home policy. - Mental and physical wellness programs. - Phone and internet reimbursement. - Continued career development support. - Comprehensive benefits and competitive compensation packages. - Paid volunteer time and employee resource groups.