Remote job
Senior Solutions Architect – Large Scale AI Inference
Job details
About this role
Role overview A senior pre-sales engineering role focused on large-scale AI inference in production environments. The position partners with leading AI-native companies, infrastructure providers, and large enterprises across EMEA to optimize inference workloads on multi-node GPU clusters. It blends hands-on technical architecture with community engagement, shaping the direction of next-generation inference systems at the intersection of AI and high-performance computing.
Responsibilities - Guide customers through the design, deployment, and tuning of large-scale inference workloads running on multi-node GPU clusters. - Architect inference pipelines for both dense and sparse Mixture-of-Experts (MoE) models, distributing compute across thousands of accelerators. - Improve efficiency across quantization (INT4/FP8), speculative decoding, disaggregated prefill/decode, KV cache management, and wide expert parallelism. - Collaborate with internal inference platform and library teams to accelerate customer outcomes and feed product improvements. - Grow the EMEA inference developer community through technical workshops, hackathons, and reference architectures.
Requirements - MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience. - 5+ years of experience optimizing neural network inference in production. - Deep understanding of transformer inference optimization, including quantization, disaggregated inference, speculative decoding, continuous batching, and KV cache techniques. - Practical experience with MoE inference at scale, including expert parallelism, wide expert parallelism, all-to-all communication, routing overhead, and load balancing. - Ability to engage with ML engineers, researchers, and systems architects at a deep technical level.
Nice to have - Hands-on experience with disaggregated inference tooling (for example Dynamo, NIXL, or Grove). - Understanding of GPU memory hierarchies and high-speed interconnects such as NVLink, InfiniBand, RDMA, and UCX. - Background at an advanced AI lab or large-scale AI infrastructure provider running inference on thousands of GPUs. - Published work or benchmarks in large-scale AI inference.
Benefits and work setup - Highly competitive compensation; for Poland-based candidates, the base salary range is 292,500 PLN to 507,000 PLN, determined by location, experience, and internal pay parity. - Comprehensive benefits package. - Employer is an equal opportunity employer committed to a diverse workforce and does not discriminate on protected characteristics in hiring or promotion.