Remote job
Data Scientist II
Job details
About this role
Role overview This role focuses on applying data science and machine learning to healthcare claims, authorization, and clinical datasets to surface actionable insights for utilization management. The work spans trend analysis, LLM development, anomaly detection, and the design of data pipelines and dashboards that integrate analytics into business workflows. It is a fully remote position based anywhere in the United States, with occasional travel for onboarding, team meetings, and company events.
Responsibilities - Gather business requirements and apply trend analysis, regression, and outlier detection to interpret patterns in healthcare claims and authorization data. - Apply large language models and stochastic optimization methods to support decision-making, including training, fine-tuning, and deploying models for anomaly detection, classification, and automated analysis. - Monitor industry trends, regulatory changes, and emerging fraud schemes to inform detection strategy development, maintain methodology documentation, and present findings to leadership. - Build scalable data pipelines on AWS using Airflow, Python, PySpark, and S3 to ingest, integrate, and transform large healthcare datasets, with downstream modeling in SageMaker. - Develop data models, algorithms, and simulations to assess expected impact and return on investment across medical expense, administrative cost, and clinical outcomes. - Build Tableau visualizations and partner with product owners and stakeholders to embed data science outputs into enterprise business workflows. - Perform data quality analysis to identify gaps or inconsistencies across EMR, EHR, claims, pharmacy, and Social Determinants of Health datasets.
Requirements - Master's degree in Data Science, Statistics, Biostatistics, or a related field, with 36 months of experience as an analyst or in a related role analyzing datasets. - 3 years of experience analyzing healthcare datasets including EMR or EHR, medical and pharmacy claims, and Social Determinants of Health data. - 3 years building scalable AWS Airflow pipelines with Python and PySpark on S3 and accelerating decisions through SageMaker. - 3 years applying regression, outlier detection, and other trend-identification methods in Python, Spark, and SQL. - 3 years building and implementing models, creating algorithms, and running simulations. - 3 years developing Tableau visualizations and partnering with stakeholders to integrate data science into business workflows. - 3 years developing NLP, including extractive and generative LLMs, fine-tuning, and broader LLM model development.
Benefits and work setup - Salary range of $137,000 to $161,000 per year for a 40-hour work week. - Fully remote within the United States, with occasional travel to company headquarters for onboarding, team meetings, and company events.