Remote job
AI Engineer/ML Engineer - Senior Developers - AI Training - Washington DC, US
Job details
About this role
Role overview Join an expert network of senior AI and machine learning engineers contributing to the training and evaluation of large language models. The work centers on applying deep technical expertise to audit model outputs, validate ML code, and supply high-quality human feedback that helps align AI systems. It is a paid, remote participant role with task-based assignments of typically up to one hour each, paying up to $80 per hour depending on the study.
Responsibilities - Review AI-generated explanations of model architectures, loss functions, and backpropagation for technical accuracy. - Validate ML-specific code such as training loops, preprocessing scripts, and evaluation notebooks for efficiency and correctness. - Provide detailed human feedback used in reinforcement learning from human feedback (RLHF) workflows to align models with intent, safety, and helpfulness. - Critically assess chain-of-thought reasoning and surface where model logic or numerical reasoning breaks down. - Run comparative benchmarks across model outputs using defined technical taxonomies and performance metrics.
Requirements - BS, MS, or PhD in Computer Science, Artificial Intelligence, Robotics, or a related quantitative field with a machine learning focus. - Professional experience building, deploying, or fine-tuning ML models in production environments. - Strong grasp of neural network architectures such as Transformers, CNNs, and RNNs, plus optimization techniques. - Hands-on experience with prompt engineering, RLHF, or retrieval-augmented generation (RAG) workflows. - Ability to audit complex model logic, detect training data contamination, and evaluate mathematical proofs behind ML algorithms. - High attention to detail for spotting hallucinations, biased outputs, or logical failures in AI-generated technical content.
Nice to have - Expert proficiency in PyTorch or TensorFlow/Keras. - Advanced Python with NumPy, Pandas, and Scikit-learn, plus Hugging Face Transformers experience. - Cloud or MLOps exposure including AWS SageMaker, Google Cloud Vertex AI, Weights & Biases, or LangChain. - Familiarity with vector databases such as Pinecone, Milvus, or Weaviate for RAG evaluation.
Benefits and work setup Remote, flexible task-based work with up to $80 per hour for studies targeting this expertise. Onboarding begins with a 10- to 15-minute skills assessment before participants are added to the expert pool.