Remote job
Senior Applied AI/ML Scientist
Job details
About this role
Role overview
This senior applied AI role centers on advancing agent-based capabilities inside a customer-facing product platform that processes data at significant scale. The scientist owns the quality and measurement layer of production AI, defining what "good" means for agents, designing the evaluations that measure it, and ensuring that bar holds as systems expand. The work blends model judgment, experimentation, and cross-functional partnership to translate evidence into product decisions.
Responsibilities
- Drive end-to-end development of agent and agentic features, including defining success criteria, guardrails, and evaluation methods alongside engineering, design, and product partners - Lead model selection and validation, rigorously comparing approaches and recommending adoptions grounded in measurable evidence - Design and run controlled experiments, including A/B tests and other statistical methods, to quantify the real customer impact of model changes - Establish and evolve the quality bar across multiple agent systems, partnering with engineering on evaluation infrastructure that scales beyond a single workflow - Identify improvements to core capabilities such as prompt iteration, reusable automated judges, context-window management, tool use, and adoption of new state-of-the-art models - Communicate AI capabilities, limitations, and trade-offs to technical and leadership audiences so decisions rest on evidence rather than intuition
Requirements
- Five or more years of applied AI/ML experience with a track record of shipping systems where evaluation and quality drove measurable impact - Hands-on experience designing evaluation frameworks and building or calibrating automated judges for ML, LLM, or agentic systems - Strong judgment in comparing models or configurations, with the ability to articulate why one approach outperforms another - Statistical rigor in experimentation, including A/B testing, power analysis, and sound measurement design - Working knowledge of Python, common ML frameworks, transformers, and embeddings, plus SQL fluency with large datasets - Ability to collaborate with product, design, and engineering and translate technical quality questions into actionable stakeholder guidance
Nice to have
- Direct experience evaluating production LLM-driven or agentic products, including prompt quality, hallucination measurement, and tool-calling reliability - Track record of defining a quality bar or evaluation methodology that other teams later built against - Rapid prototyping with emerging AI/ML techniques to assess applicability - Familiarity with evaluation and observability tooling and modern MLOps practices - Experience shaping product roadmaps with evidence about AI quality and feasibility - Awareness of model security, bias mitigation, and responsible-AI practices
Benefits and work setup
- Fully remote position; candidates must reside in Alberta, British Columbia, or Ontario - 100% employer-paid premiums for medical, dental (basic and major), vision, life insurance, and disability coverage for employees and eligible dependents - 25 days of vacation, 5 paid sick days, public holidays, and additional company-wide rest days - Comprehensive mental health support including coaching, therapy sessions, and digital wellness resources for employees and dependents - Annual lifestyle spending account, a one-time home office stipend, and a monthly internet allowance - No-cost financial wellness coaching, subsidized child and eldercare access, and charitable donation matching - Disclosed hiring range of CA$140,700 to CA$177,900 annually, within a broader band of CA$140,700 to CA$211,100, with pay determined by location, experience, and skills