Remote job
Solutions Architect - Deep Neural Network Evaluation
Job details
About this role
Role overview
Serve as a customer-facing AI solutions architect focused on evaluating deep neural networks and end-to-end agentic systems. The work combines technical discovery, proof-of-concept development, benchmarking, failure analysis, and collaboration with engineering, product, sales, and business stakeholders across EMEA. The main technical scope includes large language models, vision-language models, embedding models, retrieval architectures, and production-oriented agent pipelines.
Responsibilities
- Work with customers and AI application teams to understand technical goals and design suitable evaluation and deployment approaches. - Build and demonstrate solutions using open-source and commercial LLM technologies within retrieval and agentic workflows. - Benchmark models and pipelines across varied use cases and languages, assessing performance from individual models through complete systems. - Investigate system failures and recommend mitigations such as model selection, retrieval optimization, fine-tuning, or pipeline changes. - Promote and contribute to reusable evaluation tools, measurement methods, benchmarks, and robustness practices. - Translate customer feedback and proof-of-concept findings into product improvements and practical enterprise architecture guidance.
Requirements
- Master’s or doctoral degree, or equivalent experience, in computer science, data science, engineering, physics, mathematics, or a related discipline. - At least five years of hands-on experience developing or evaluating deep neural networks. - Professional or academic experience in machine learning, deep learning, or data science, with strong familiarity with current LLMs, VLMs, and retrieval systems. - Knowledge of model-evaluation libraries and services, agentic pipeline assessment, and modern benchmark design; ability to extend existing benchmarks or create new ones. - Strong written, verbal, and technical presentation skills in English. - Ability to collaborate effectively with customers, engineering groups, product teams, sales, and other business stakeholders.
Nice to have
- Experience assessing memorization, hidden-instruction risks, security concerns, or other robustness properties of models. - Experience evaluating large models at scale, understanding distributed training, or applying evaluation methods to reinforcement learning.