← Back to jobs

Remote job

Senior Solutions Architect – Large Scale AI Inference

Other Full-time Permanent France

Job details

Not specified Salary
France Eligibility
Senior Experience
Full-time Employment

About this role

Role overview A senior pre-sales engineering role focused on large-scale AI inference in production environments. The position partners with leading AI-native companies, infrastructure providers, and large enterprises across EMEA to optimize inference workloads on multi-node GPU clusters. It blends hands-on technical architecture with community engagement, shaping the direction of next-generation inference systems at the intersection of AI and high-performance computing.

Responsibilities - Guide customers through the design, deployment, and tuning of large-scale inference workloads running on multi-node GPU clusters. - Architect inference pipelines for both dense and sparse Mixture-of-Experts (MoE) models, distributing compute across thousands of accelerators. - Improve efficiency across quantization (INT4/FP8), speculative decoding, disaggregated prefill/decode, KV cache management, and wide expert parallelism. - Collaborate with internal inference platform and library teams to accelerate customer outcomes and feed product improvements. - Grow the EMEA inference developer community through technical workshops, hackathons, and reference architectures.

Requirements - MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience. - 5+ years of experience optimizing neural network inference in production. - Deep understanding of transformer inference optimization, including quantization, disaggregated inference, speculative decoding, continuous batching, and KV cache techniques. - Practical experience with MoE inference at scale, including expert parallelism, wide expert parallelism, all-to-all communication, routing overhead, and load balancing. - Ability to engage with ML engineers, researchers, and systems architects at a deep technical level.

Nice to have - Hands-on experience with disaggregated inference tooling (for example Dynamo, NIXL, or Grove). - Understanding of GPU memory hierarchies and high-speed interconnects such as NVLink, InfiniBand, RDMA, and UCX. - Background at an advanced AI lab or large-scale AI infrastructure provider running inference on thousands of GPUs. - Published work or benchmarks in large-scale AI inference.

Benefits and work setup - Highly competitive compensation; for Poland-based candidates, the base salary range is 292,500 PLN to 507,000 PLN, determined by location, experience, and internal pay parity. - Comprehensive benefits package. - Employer is an equal opportunity employer committed to a diverse workforce and does not discriminate on protected characteristics in hiring or promotion.

Skills detected in the listing

LLM
Detected Sep 8, 2026
Last verified Sep 8, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight