Remote job
Senior Data & LLM Engineer (AI-Ready Data & Agentic Systems)
Job details
About this role
Role overview
A boutique consultancy is hiring a senior engineer for a six-month embedded engagement inside a private-capital data team. The seat deliberately blends two halves: senior data engineering on a modern cloud stack, and senior LLM/agentic engineering focused on systems already in production. Day-to-day work centers on two initiatives — building an AI-ready data catalog and context layer alongside the analytics team, and scaling unstructured data extraction — with agentic implementation, evaluation, and orchestration as the connective tissue across both. This is not a maintenance role; the expectation is someone who has shipped agents into production and can speak candidly about what broke.
Responsibilities
- Design and build production agents and multi-step agentic workflows, including custom MCP servers and tool interfaces that expose proprietary datasets to agents. - Stand up an evaluation practice from scratch — eval datasets, LLM-as-judge, regression suites — and track task success, faithfulness, extraction accuracy, tool-call correctness, latency, and cost per run. - Orchestrate agentic workloads in Airflow 3.x on Astronomer, modeling LLM and agent calls as named, independently retriable tasks with dynamic fan-out/fan-in and human-in-the-loop operators. - Decide what belongs outside the DAG (long-running agents on AWS ECS/Fargate, queue- and event-driven execution, API-triggered services) and operate those services. - Engineer AI-ready datasets, embeddings, and retrieval pipelines, treating dbt metadata as a first-class input to the context layer. - Build operational hardening for non-deterministic systems: idempotency, retry semantics, budget and token caps, circuit breakers, structured logging, and audit trails under Git-based CI/CD.
Requirements
- 7+ years of hands-on industry experience in data or software engineering, with at least 5 years on a major cloud platform; bachelor's degree or higher in a technical field, or equivalent experience. - Demonstrated track record shipping LLM-powered or agentic systems to production, with concrete examples of what was built, how it was evaluated, and how it failed. - Practical command of agent design patterns: tool and function calling, structured outputs, retrieval/RAG, planning loops, multi-agent decomposition, and human-in-the-loop. - Hands-on discipline with agent evaluation and observability — eval datasets, LLM-as-judge, regression testing, tracing, and cost/latency monitoring. - Daily fluency with a frontier model API such as Claude or comparable, plus an agent framework (Claude Agent SDK and/or OpenAI Agents SDK / Assistants API) and a clear point of view on framework vs. custom loop. - Experience connecting models to enterprise data through MCP and/or comparable tool-integration patterns, plus strong Snowflake, dbt, Airflow, AWS, and Python ingestion skills.
Nice to have
- Hands-on experience with Snowflake Cortex (Analyst, Search, Agents, or AISQL). - Semantic layer or data catalog exposure (dbt Semantic Layer, Cube, AtScale, DataHub, OpenMetadata, Atlan, Collibra, or Snowflake Horizon). - Knowledge graph or ontology modeling, or prior work with `dlt` and custom connectors. - Streaming or event-driven experience (Kafka, Kinesis, SNS/SQS). - MLOps depth beyond agents, LLM fine-tuning and deployment, or publications in relevant AI/ML communities. - Background in private equity, venture capital, or financial services data environments, and prior embedded-consulting experience.
Benefits and work setup
- Fully remote with a flexible schedule. - Unlimited paid time off, plus paid parental and bereavement leave. - Eastern Time business hours (09:00–17:00 ET) overlap expected. - Engagement with globally recognized clients and collaboration with a senior technical team.