Remote job
Code Data Specialist
Job details
About this role
Role overview This position supports the development of frontier AI models by producing high-quality training data across a range of tasks. The work spans reinforcement learning from human feedback (RLHF), instruction tuning, and preference labeling, all aimed at strengthening how large language models interpret and respond to user input.
Responsibilities - Apply RLHF labeling practices to shape model behavior through human feedback signals - Build and curate training examples designed to improve instruction-following capability - Assess and rank model outputs to inform preference optimization efforts - Compose detailed prompts and craft high-quality responses spanning multiple subject areas - Record edge cases and document failure modes observed in model behavior for ongoing improvement
Requirements - Demonstrated analytical and problem-solving ability - Strong written communication skills in English, with the ability to produce clear, accurate annotations - Working familiarity with large language models and core prompt engineering concepts - Careful attention to detail and the discipline to follow structured annotation guidelines - Prior exposure to data labeling or content moderation workflows
Nice to have - Hands-on experience contributing to RLHF datasets or similar human-feedback pipelines - Background in evaluating generative AI outputs against quality rubrics