Remote job
Evaluations Team Lead
Job details
About this role
Role overview
A player-coach leadership role building the evaluation function for a production-grade tabular foundation model. The team will own how the model is measured, creating the shared benchmark platform that research, engineering, and customer-facing groups all rely on, and setting the methodology the rest of the company uses to judge quality, latency, and cost.
Responsibilities
- Establish and operate a shared evaluation platform with curated datasets, a model registry, and reproducible pipelines that let teams run benchmarks and investigate results independently. - Define and uphold evaluation standards covering metrics, data splits, leakage prevention, calibration, uncertainty quantification, and benchmark contamination. - Integrate regression checks into research and release workflows, using versioned inputs, versioned artifacts, and end-to-end latency measurement. - Run fair head-to-head comparisons against competing approaches, matching tuning budgets, data access, compute, and latency protocols. - Support customer proof-of-concept engagements with tooling, methodological guidance, and analysis. - Measure predictive quality, latency, and cost across deployment configurations, task types, and dataset characteristics, then translate findings into research priorities, release recommendations, and evidence-backed improvement plans. - Hire, develop, and prioritize work for a small team while remaining hands-on with code, experiment design, and statistical methodology.
Requirements
- Demonstrated experience owning evaluation for production tabular ML systems that drive consequential decisions. - Strong statistical judgment across metric selection, validation schemes, uncertainty estimation, cross-dataset model comparison, and accounting for repeated experimentation. - Practical track record finding leakage in preprocessing, feature construction, joins, temporal dependencies, and related-entity splits. - Strong Python and SQL fluency, working knowledge of scikit-learn and gradient-boosted trees, and a history of building reliable ML tooling used by other teams. - Experience designing fair model comparisons, including hyperparameter search, resource budgets, and end-to-end latency measurement. - Prior people management that included hiring, technical coaching, and performance feedback while remaining technically involved. - Clear written and verbal communication with researchers, engineers, and customer data scientists, including the willingness to push back on claims the evidence does not support.
Nice to have
- Experience evaluating tabular foundation models or AutoML systems. - Experience measuring how optimization choices trade off predictive quality against inference performance. - Familiarity with relational or multi-table data, and with Snowflake or Databricks environments. - Experience evaluating automated or agent-driven ML workflows, including failure modes that aggregate metrics can hide.
Benefits and work setup
- Competitive compensation comprising salary and equity. - Comprehensive health coverage for the employee and dependents. - Paid parental leave for all new parents, including adoptive and surrogate journeys. - Relocation support for employees joining in designated office hubs. - Mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action.