Remote job
Senior Solutions Architect, Customer Success and Partnership
Job details
About this role
Role overview
Serve as a senior technical partner for customers and internal teams delivering large-scale networking and AI infrastructure projects. The role combines hands-on analysis of GPU-accelerated systems with architecture guidance, performance optimization, automation, operational resilience, and customer-facing technical leadership. Work may involve customer engagements and travel of up to 25%.
Responsibilities
- Analyze, tune, and optimize complex GPU-accelerated systems and AI workloads across customer data-center environments. - Act as a senior technical authority during architecture reviews and infrastructure decisions involving networking, systems, automation, and large-scale operations. - Develop monitoring, telemetry, analytics, and optimization methods to identify bottlenecks, improve availability, and strengthen infrastructure resilience. - Lead technical projects from initial design through implementation, post-deployment improvement, SLA alignment, and risk mitigation. - Participate in incident retrospectives and post-deployment reviews, turning operational findings into customer guidance and infrastructure improvements. - Identify infrastructure opportunities in cloud and enterprise settings and support technical initiatives that demonstrate practical AI-platform value. - Coordinate with customers, partners, and internal engineering groups while guiding discussions, influencing decisions, and maintaining strong working relationships.
Requirements
- At least 10 years of experience operating large-scale data-center services, with a focus on infrastructure. - Bachelor’s, master’s, or doctoral degree in computer science, electrical or computer engineering, physics, mathematics, or a related field, or equivalent experience. - Strong analytical, troubleshooting, decision-making, time-management, organizational, and communication skills. - System-level expertise spanning operating systems, Linux kernel drivers, GPUs, network interface hardware, and computer architecture. - Experience with cloud orchestration and workload scheduling platforms, including Kubernetes, Docker Swarm, or HPC schedulers such as Slurm. - Familiarity with cloud-native technologies and their integration with traditional infrastructure. - Ability to lead complex projects, determine root causes, guide technical teams, and deliver resilient solutions.
Nice to have
Deep familiarity with AI infrastructure and training or inference workflows; MLOps and DevOps tools; containerized deployments; large-scale systems; data-center safety, security, environmental controls, and operating procedures; and certifications in data-center, server, or networking technologies. Strong customer-facing collaboration and willingness to travel are also valuable.