Remote job
Senior Data Scientist
Job details
About this role
Role overview Own the development and reliability of supervised models that produce property-related estimates and scores. The work spans defining what a model should predict, building features from complex data, validating performance in realistic settings, and explaining where results can and cannot be trusted.
Responsibilities - Train and maintain models on large tabular datasets with substantial geographic information, comparing gradient-boosted trees with linear and generalized linear baselines. - Define prediction targets and assess whether available data can support them responsibly. - Build features from location, property structure and condition, time, and linked public records. - Design validation that holds out both locations and time periods to reduce leakage and test real-world generalization. - Calibrate predictions, monitor model drift, and decide when a model should be retired. - Document training data, limitations, and appropriate use, then communicate results so non-specialists can make informed decisions.
Requirements - At least three years building supervised models on tabular data used by production systems. - Strong Python skills, including pandas, NumPy, and scikit-learn, plus practical knowledge of a gradient-boosting library and its failure modes. - Ability to write and troubleshoot SQL queries against large databases, including investigating slow query plans. - Experience handling incomplete fields, duplicate records, noisy labels, and conflicting data sources. - Willingness and ability to identify model errors, explain limitations, and raise concerns proactively.