Remote job
AI Engineer
Job details
About this role
Role overview Build and operate production-grade LLM and agent systems for real client workflows. This role covers the full path from prototyping to reliable deployment, with particular ownership of evaluation, safety controls, observability, cost management, permissions, and human oversight.
Responsibilities - Design and deliver AI and agent features from initial concept through production release. - Create evaluation suites that measure system quality and track performance across releases. - Implement guardrails, access controls, cost controls, monitoring, and escalation paths around model behavior. - Investigate the difference between test performance and real-world production outcomes using structured instrumentation and analysis. - Work directly with client stakeholders to understand operational workflows, define the right technical scope, and deliver useful systems. - Help establish dependable human-in-the-loop processes for decisions that require review or intervention.
Requirements - Practical experience building and shipping LLM-powered or agent-based systems. - Ability to turn qualitative AI behavior into measurable evaluation criteria and repeatable tests. - Strong debugging and production-observability skills, including the ability to identify gaps between controlled evaluations and live usage. - Sound judgment around model reliability, permissions, safety, cost, and human review. - Ability to communicate clearly with both technical teams and client stakeholders. - A hands-on, end-to-end approach to taking AI systems from prototype to dependable operation.
Benefits and work setup The role is remote-friendly and async-first, supporting work with clients across different regions. Candidates can expect to work closely with the people building and using the systems, with hiring emphasis placed on demonstrated work and specific, substantive applications.