Remote job
Staff Engineer / AI Builder
Job details
About this role
Role overview
This Staff AI Builder position leads the design and delivery of production AI/ML systems for an engineering team focused on enterprise-grade work. The role suits an experienced engineer who can own systems end-to-end, make architectural calls confidently, and contribute to technical leadership across the team. It centers on agentic and generative AI systems where reliability, security, and scalability are critical to real-world deployments, with an emphasis on shipping production-ready work on tight timelines.
Responsibilities
- Design agentic workflows including single- and multi-agent architectures, reasoning loops, tool/function calling, and orchestration patterns chosen for the problem at hand - Build and maintain retrieval-augmented generation pipelines covering chunking, embeddings, vector search, re-ranking, and refresh handling for dynamic knowledge sources - Integrate with GenAI services and agent frameworks, including MCP-based tool design with descriptions precise enough to drive correct routing - Write production-grade system prompts with structured role framing, output constraints, and few-shot design - Build evaluation and observability into shipped agents using golden datasets, RAGAS-style metrics, LLM-as-judge as both runtime guardrail and offline eval, and tracing via tools like LangFuse or LangSmith - Design backend services in Python and Node.js, including serverless architectures, single-table NoSQL schemas for conversation state and memory, and event-driven orchestration - Mentor engineers and contribute to technical leadership across cross-functional teams, owning features and releases end-to-end including debugging and hardening
Requirements
- 6+ years of professional software engineering experience with meaningful shipped GenAI/LLM-powered systems in production, not just prototypes - Hands-on depth in agentic AI: reasoning loops, tool/function calling, multi-agent orchestration, with a real perspective on when single-agent design beats multi-agent - Practical RAG expertise covering chunking strategies, embeddings, vector databases, cosine similarity search, and re-ranking - Experience building evaluation and observability for LLM systems, including golden datasets, LLM-as-judge, RAGAS or comparable metrics, and tracing tooling - Strong prompt engineering skills with the ability to write a production prompt with real structure and constraints on request - Hands-on experience with the AWS GenAI stack including Bedrock, Lambda, DynamoDB single-table design, S3, SQS, EventBridge, and Step Functions - Strong Python and Node.js skills with full-stack application and RESTful API experience - Solid grasp of production reliability patterns for LLM-backed systems: retries with backoff and jitter, circuit breakers, and fallback models - Experience with containerization (Docker) and cloud-native deployment - Comfort owning ambiguous, integration-heavy problems and going a level deeper when challenged
Nice to have
- Direct experience with Bedrock Agent Core, including agent registration, MCP tool wiring, memory/session management, and gateway provisioning - Background in a regulated or high-stakes domain such as education, healthcare, or financial services where evaluation rigor carries real consequences - Familiarity with workflow orchestration tools such as Step Functions or similar state-machine-based systems - A point of view on LLM observability tooling gained from running it in production
Benefits and work setup
- Salary range of $176,612 - $243,680 CAD