Remote job
AI Engineer/ML Engineer - Senior Developers - AI Training - Boston, US
Job details
About this role
Role overview
Join an expert network of machine learning professionals contributing to the training and evaluation of large language models. This is a participant-based, paid contract role based in Boston where your technical expertise directly shapes AI model accuracy, safety, and reasoning quality. Successful candidates complete paid study tasks, typically requiring up to one hour of focused work, on a flexible remote schedule.
Responsibilities
- Review AI-generated explanations of model architectures, loss functions, and backpropagation for technical correctness - Audit machine learning code, training loops, preprocessing scripts, and evaluation notebooks for efficiency and accuracy - Provide structured human feedback that supports reinforcement learning from human feedback (RLHF) workflows and model alignment - Critically assess how models navigate complex chain-of-thought prompts and pinpoint where reasoning breaks down - Benchmark and compare model outputs against defined technical taxonomies and performance metrics
Requirements
- BS, MS, or PhD in Computer Science, Artificial Intelligence, Robotics, or a related quantitative field with a machine learning focus - Hands-on production experience building, deploying, or fine-tuning ML models - Professional-level understanding of neural network architectures such as Transformers, CNNs, and RNNs, plus optimization techniques - Practical experience with prompt engineering, RLHF, or retrieval-augmented generation (RAG) workflows - Ability to audit model logic, flag potential training-data contamination, and evaluate the mathematical proofs behind ML algorithms - Strong attention to detail for spotting hallucinations, biased outputs, or logical failures in AI-generated technical content
Nice to have
- Proficiency with PyTorch or TensorFlow/Keras - Advanced Python skills including NumPy, Pandas, Scikit-learn, and Hugging Face Transformers - Experience with cloud MLOps platforms such as AWS SageMaker, Google Cloud Vertex AI, Weights & Biases, or LangChain - Familiarity with vector databases like Pinecone, Milvus, or Weaviate for RAG evaluation
Benefits and work setup
- Hourly compensation commonly up to $80, depending on study requirements - Fully remote with flexible scheduling aligned to available studies - Short 10–15 minute skills assessment as the entry point, with quick onboarding once accepted