Remote job
Senior Solutions Architect – Large Scale Neural Networks Inference
Job details
About this role
Role overview This senior role defines the technical direction for large-scale AI inference across the EMEA region, partnering with frontier AI labs and enterprises that deploy AI systems at scale. The position blends deep expertise in neural network optimization with strategic leadership, identifying performance bottlenecks and driving high-impact solutions that shape next-generation inference deployments. It combines hands-on architecture work with cross-organizational alignment to influence product roadmaps.
Responsibilities - Lead inference strategy for a portfolio of EMEA AI-native customers, guiding engagements from initial proof of concept through production-scale deployment - Diagnose inference challenges across customer deployments, spanning latency, efficiency, cost per token, memory utilization, and low-latency networking - Architect and optimize high-performance inference pipelines using modern inference backends, improving GPU utilization and AI cluster efficiency - Translate customer insights and deployment patterns into actionable product feedback that informs the inference platform roadmap
Requirements - MS or PhD in Computer Science, Engineering, or equivalent practical experience - 8+ years in AI/ML infrastructure experience with deep expertise in LLM/VLM inference optimization and production-scale deployment - Strong command of transformer inference acceleration, including quantization (INT4/FP8), speculative decoding, disaggregated inference, continuous batching, KV cache optimization, and wide expert parallelism for MoE models - Understanding of GPU memory hierarchies and low-latency networking and their influence on inference performance - Proven track record leading technical initiatives, with strong communication skills across research scientists, infrastructure engineers, and executive audiences
Nice to have - Experience with proprietary GPU inference stacks and serving frameworks - Background orchestrating GPUs on Kubernetes - Hands-on experience operating inference at scale inside a frontier AI lab or hyperscale inference team - Contributions to open-source inference projects such as vLLM, SGLang, or KServe
Benefits and work setup Compensation is highly competitive and accompanied by a comprehensive benefits package. The organization is an equal opportunity employer committed to a diverse workforce and does not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law. Base pay is determined by location, experience, and the pay of employees in similar positions.