Remote job
AI Engineer – Trust & Explainability (AI Platform)
Job details
About this role
Role overview L2 AI Engineer role on an AI Platform team building the shared runtime that powers every AI agent across a multi-tenant SaaS product for credit unions and lenders. The work focuses on the trust and explainability layer: tracing, evaluation, and human-readable explanations for multi-agent workflows, with a dotted-line relationship to the Security Operations team to align threat models and platform guarantees.
Responsibilities - Build tracing across the gateway, orchestration, memory, and tool layers of the AI runtime, including correlation across handoffs, parallel branches, and retries in a multi-agent workflow. - Develop the developer-facing trace view and the explanation layer that turns trace data into a readable account of what an agent did and why. - Ship the platform primitives product teams surface to end users: explanation records, confidence and provenance metadata, and summaries of what the agent relied on. - Evaluate and integrate open-source observability, tracing, and evaluation frameworks, extending them where agent workloads need more. - Build CI evaluation checks using golden datasets, rubric and LLM-as-judge scoring, and regression baselines for model, prompt, or tool changes. - Build red-team and adversarial test suites covering prompt injection, jailbreaks, tool misuse, and data exfiltration, plus automated tests proving agents cannot cross tenant boundaries. - Participate in code review, mentor L1 AI Engineer teammates, and contribute documentation that reduces tribal knowledge.
Requirements - 3+ years of professional software engineering experience delivering features independently in production. - Hands-on experience building software that integrates LLMs (agent frameworks, RAG pipelines, or similar), in production or substantial projects. - Proficiency in Python with production experience; TypeScript is a plus. - Solid grasp of algorithms, data structures, and modern development tooling including Git, Docker, and automated testing. - Hands-on experience in at least one of: LLM observability and tracing, LLM evaluation and testing, agent frameworks and multi-agent orchestration, or application security testing. - Experience with distributed tracing or observability tooling in a production system, plus running workloads on Azure or AWS (identity, networking, secrets management). - Active daily use of AI-assisted development tools and a Bachelor's degree in Computer Science or equivalent experience.
Nice to have - Familiarity with OpenTelemetry, including GenAI semantic conventions, or OpenLLMetry. - Experience with LLM observability and evaluation tools such as Langfuse, Arize Phoenix, LangSmith, Braintrust, promptfoo, or DeepEval. - Contributions to open-source AI observability, evaluation, or agent framework projects. - Experience with AWS Bedrock, Azure OpenAI, or other cloud-managed model services. - Background operating multi-tenant SaaS systems where tenant isolation was a hard requirement. - Prior work in financial services, fintech, or another regulated industry where explainability shaped technical decisions. - Experience building developer-facing debugging or visualization tools.