Remote job
Senior AI Engineer
Job details
About this role
Role overview A full-time role on an agentic AI team building production-grade AI systems that support U.S. government and industrial customers. The position sits at the intersection of applied AI and systems engineering, focused on agent architectures, model integrations, evaluation systems, and infrastructure for moving agentic AI beyond prototypes into reliable, real-world software.
Responsibilities - Design and build production agentic AI systems that handle reasoning, planning, tool use, context management, memory, and multi-step task execution - Develop infrastructure supporting multiple commercial and open-weight language models, including inference, serving, and scaling concerns - Build automated evaluation frameworks that measure agent quality, reliability, task completion, and regressions - Create datasets, benchmarks, and evaluation methodologies for complex agentic workflows - Improve observability through structured logging, metrics, distributed tracing, dashboards, and automated alerting - Debug failures spanning model behavior, agent execution, application code, and distributed infrastructure, and turn findings into systematic improvements
Requirements - U.S. citizenship is required - 5+ years of experience building production software, AI/ML systems, or distributed systems - Deep experience designing, building, and operating production AI agents or agent platforms - Strong Python programming skills and modern LLM knowledge, including inference, context engineering, tool calling, retrieval, structured outputs, and model selection - Hands-on experience with Kubernetes and a major cloud platform such as AWS, GCP, or Azure - Comfort moving between research and engineering: prototype an idea, evaluate it rigorously, and turn successful approaches into production systems
Nice to have - Current or obtainable U.S. security clearance, with sponsorship available - Background in startups or other fast-paced entrepreneurial environments - Experience with multi-agent architectures, secure code execution sandboxes, or AI observability tooling - Exposure to fine-tuning, post-training, reinforcement learning, or synthetic data generation - Work in defense, government, or other mission-critical environments
Benefits and work setup - Based in Pittsburgh, PA or remote; up to 25% travel for team collaboration and planning