Remote job
Staff AI Platform Engineer: Agent & Retrieval Infrastructure
Job details
About this role
Role overview
Lead the architecture and delivery of a production AI platform that supports autonomous agents, retrieval, and secure access to operational and scientific data. This staff-level role combines backend engineering, data engineering, cloud infrastructure, security, and platform enablement, with responsibility for turning managed foundation-model services into reliable systems used by other engineers and, eventually, external users.
Responsibilities
- Design the agent platform, including orchestration, action groups, backend APIs, model access, deployment strategy, and environment separation. - Own the retrieval data plane from ingestion and document chunking through embeddings, indexing, storage, freshness, and search-quality optimization. - Extend pipelines for heterogeneous internal, geospatial, scientific, and large-binary datasets, building missing components where necessary. - Establish security boundaries with guardrails, private networking, least-privilege machine identities, encryption, audit trails, and controlled tool access. - Define governance for autonomous actions, including approval thresholds, monitoring, incident response, and operational limits. - Build observability and evaluation systems for traces, tool calls, retrieval performance, model quality, release gates, and ongoing regression detection. - Provide reusable abstractions, SDKs, infrastructure-as-code, and self-service environments so engineers can deliver AI features independently.
Requirements
- At least eight years of software and infrastructure engineering experience, including deep production backend work and staff-level technical ownership. - Strong hands-on experience deploying managed agent and retrieval services in production, including model access, knowledge bases, guardrails, throughput, and quota management. - Experience operating containerized or serverless workloads and owning CI/CD, deployment safety, and production operations. - Practical expertise with RAG, embeddings, chunking strategies, semantic search, and production vector stores such as managed OpenSearch, Pinecone, or pgvector. - A track record building ingestion systems over messy, unstructured data while treating freshness, correctness, and availability as operational commitments. - Deep cloud infrastructure knowledge covering identity and access management, private networking, object storage, key management, monitoring, and Terraform, CDK, or CloudFormation. - Production exposure to LLM features or autonomous agents, with sound judgment about access controls, failure modes, observability, and secure rollout. - Strong Python or TypeScript skills; Go experience is also relevant. Comfortable working across a small team and documenting systems for independent operation.
Nice to have
- Experience with automated LLM evaluations, release gates, AI red-teaming, knowledge graphs or GraphRAG, geospatial or scientific data, intermittent connectivity, compliance programs, or autonomous and field operations.