Job details
About this role
Role overview Help train language models to tell the difference between depictions of violence and content that enables it, evaluating user requests, model responses, and conversation context against customer policy. This is a reasoning-heavy policy role that requires holding a steady line through ambiguous, emotionally heavy material.
Responsibilities - Evaluate user requests and AI model responses involving violence, weapons, threats, and dark fiction within full conversation context - Distinguish fictional, educational, historical, and defensive violence from requests seeking real-world uplift or expressing genuine intent to harm - Assess whether a model response provides meaningful real-world capability regardless of how the request was framed - Tell apart expressions of anger, frustration, or dark humor from credible threats or crisis indicators - Choose the most defensible classification on genuinely ambiguous cases and write concise rationales citing policy language and conversation details - Write and refine adversarial or borderline prompts that probe where a model draws its line - Surface policy gaps, contradictions, and emerging edge cases to project leads and policy teams - Participate actively in calibration discussions, challenge interpretations respectfully, and update judgment when stronger reasoning emerges - Apply customer policy consistently without substituting personal beliefs for the standard - Maintain accuracy across repeated, feedback-heavy evaluations
Requirements - Strong instincts about violence in at least one area such as fiction, real-world violence, or crisis behavior - Attention to the detail that shifts outcomes, since one word or one contextual cue can change the answer - Capacity to hold a strong opinion without becoming attached to being right - Ability to explain judgment calls clearly enough to be audited - Capacity to separate personal views from the customer-supplied standard - Careful, consistent work through repetitive exposure to difficult material - Clear and precise written communication
Nice to have - Published or produced work involving violence, including novels, short fiction, screenplays, comics, tabletop or video game content, mods, or audience-facing fan fiction - Military, law enforcement, corrections, private security, or armed professional experience - Training or professional experience in crisis intervention, threat assessment, forensic or clinical psychology, or violence prevention - Experience with firearms, martial arts, or other weapons disciplines as instructor, competitor, or professional - Prior AI evaluation, red teaming, data annotation, RLHF, trust and safety, or content moderation work - Professional experience evaluating outputs from ChatGPT, Claude, Gemini, or comparable language models - Familiarity with calibration sessions, inter-rater agreement, or adjudication workflows
Benefits and work setup - Remote within the US - Compensation: $45-55/hour - W-2 employment classification - Schedule: 8 AM to 5 PM Pacific Time, Monday through Friday - Structured evaluation frameworks with exposure limits, content rotation, and access to mental health support