Remote job
Senior Machine Learning Engineer
Job details
About this role
Role overview A Senior Machine Learning Engineer builds the infrastructure that powers the full lifecycle of language models and AI systems for an AI-native enterprise readiness platform serving government and industrial customers. The role owns LLMOps, fine-tuning infrastructure, evaluation, dataset pipelines, experiment management, serving, and observability. It can be based in Pittsburgh, PA, or remote, with up to 25% travel for team collaboration and in-person planning.
Responsibilities - Design and build LLMOps infrastructure supporting development, evaluation, deployment, and continuous improvement of production language models. - Build scalable training and fine-tuning infrastructure, including pipelines for supervised fine-tuning, parameter-efficient techniques, and preference optimization. - Develop data pipelines for training, fine-tuning, evaluation, and synthetic data generation, plus dataset versioning, lineage, and quality validation. - Build experiment management infrastructure for comparing models, datasets, hyperparameters, prompts, and training techniques. - Operate scalable model-serving and inference infrastructure, with abstractions that decouple applications from any single model or vendor. - Build observability for training and inference, optimize GPU utilization, latency, throughput, and cost, and automate deployment, rollback, and canarying workflows.
Requirements - U.S. Citizenship is required. - 5+ years building production machine learning systems, ML infrastructure, or distributed systems. - Deep experience with the modern LLM lifecycle, including training, fine-tuning, evaluation, deployment, inference, and monitoring. - Hands-on experience with LLM fine-tuning and post-training workflows such as supervised fine-tuning, LoRA/QLoRA, or preference optimization. - Strong Python programming and experience with containers, Kubernetes, and cloud platforms such as AWS, GCP, or Azure. - Solid understanding of distributed systems and comfort debugging failures across training code, datasets, models, GPUs, and cloud infrastructure.
Nice to have - Ability to obtain a U.S. security clearance with company sponsorship. - Experience with startup environments, secure code execution or sandboxes for AI agents, multi-agent architectures, or synthetic data generation. - Background in AI observability, tracing, and debugging infrastructure, or optimizing inference latency, throughput, GPU utilization, and model-serving costs. - Familiarity with AI security, adversarial testing, or securing agentic systems, including work in government, defense, or other mission-critical environments.
Benefits and work setup - Full-time role based in Pittsburgh, PA, or remote, with up to 25% travel, including periodic visits to Pittsburgh and Arlington, VA offices.