Remote job
Senior Solutions Architect - Multimodal AI
Job details
About this role
Role overview
A senior pre-sales technical position focused on guiding customers across the EMEA region who are building production-grade multimodal AI systems. The role sits at the intersection of applied research and field engineering, helping teams architect, optimize, and deploy solutions for document intelligence, personalization, and image or video understanding. The work blends deep technical advisory with active engagement of the external developer community.
Responsibilities
- Build trusted technical relationships with teams developing multimodal AI systems for document intelligence, personalization, and visual content analysis, supporting them from architecture planning through to production deployment. - Advise on training strategies that span layout-aware document encoders, user-item interaction models, and spatiotemporal video representations. - Tackle practical vision challenges including image resolution tradeoffs, vision encoder optimization for production latency budgets, efficient video frame sampling, and temporal reasoning. - Translate field insights into product roadmap input across model training, inference optimization, and data science acceleration tooling. - Engage the wider developer community through hackathons, technical talks, live demos, and reference architecture blueprints.
Requirements
- MS or PhD in Computer Science, Engineering, or equivalent practical experience. - Seven or more years of applied AI/ML experience with hands-on work in document understanding and visual content analysis. - Proven track record building or optimizing vision-language models and omnimodal architectures. - Familiarity with GPU-accelerated training and inference frameworks for large deep learning models. - Excellent communication skills, with the ability to bridge research scientists, ML engineers, and business stakeholders.
Nice to have
- Experience building multimodal systems that combine audio and video with complex structured content such as tables, layouts, and multi-page reasoning. - Understanding of retrieval and search systems including dense retrieval, approximate nearest neighbor indexing, and re-ranking pipelines. - Production optimization of vision encoders through quantization, pruning, or architectural changes to meet real latency targets. - Published research or open-source contributions in multimodal learning.
Benefits and work setup
Highly competitive compensation package with comprehensive benefits. Base pay is determined by location, experience, and internal benchmarks for comparable roles. The organization is committed to equal opportunity and a diverse workforce, and does not discriminate on the basis of protected characteristics in hiring or promotion.