Remote job
Machine Learning Engineer – ML Evaluation & Experiment Design
Job details
About this role
Role overview A specialized consulting project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation. Engineers will analyze ML experiments, datasets, metrics, and pipelines to determine whether each challenge is technically sound, reproducible, appropriately difficult, and genuinely rewards good ML reasoning rather than brute-force model selection or sweeping hyperparameter search. It is a strong fit for ML engineers who enjoy debugging experiments and designing rigorous evaluations.
Responsibilities - Review ML challenges to determine whether they are well designed and technically solvable. - Evaluate whether datasets contain meaningful, learnable signals or rely on artifacts and shortcuts. - Detect metric gaming, data leakage, label noise, distribution shift, and other evaluation flaws. - Identify unintended shortcuts or spurious correlations in synthetic datasets. - Verify reproducibility across the complete data → model → evaluation pipeline. - Assess whether challenge difficulty is appropriately calibrated and provide clear recommendations to recalibrate or exclude problematic tasks.
Requirements - 3+ years of hands-on applied machine learning experience. - Strong experience with experiment design, model selection, hyperparameter tuning, and model evaluation. - Deep understanding of train, validation, and test splits and sound preprocessing practice. - Ability to identify data leakage, label noise, distribution shift, spurious correlations, and contamination. - Familiarity with statistical significance, effect sizes, and choosing appropriate ML evaluation metrics. - Experience debugging ML workloads across CPU and GPU environments.
Nice to have - Experience creating or participating in Kaggle, DrivenData, or similar ML competitions. - Background in data-centric AI, dataset quality, or synthetic data generation and validation. - Familiarity with statistical testing, confidence intervals, and effect sizes. - Experience with ML evaluation pipelines, RLHF, or AI model evaluation more broadly. - Understanding of common ML failure modes such as shortcut learning, Goodhart's Law, Simpson's paradox, and metric gaming.
Benefits and work setup Remote, part-time, project-based consulting engagement focused on applied machine learning, experiment design, data quality, and model evaluation.