Job details
About this role
Role overview Help train image generation models by evaluating prompts and outputs for adherence, quality, and content policy compliance. This is a critical-judgment role where evaluators reason through close calls, write defensible rationales, and feed that reasoning back into model training pipelines.
Responsibilities - Evaluate generated images against their prompts for adherence, composition, realism, style consistency, and technical defects - Compare images side by side and select the stronger one with clear, evidence-based reasoning - Classify images against customer content policies covering sexual content, violence, hate symbols, real-person likeness, intellectual property, and depictions of minors - Make defensible calls on genuinely ambiguous cases and write concise rationales citing rubric language and specific visual details - Separate personal taste from prompt failure from policy violation - Write and refine probes that expose where a model's quality or safety behavior breaks down - Flag rubric gaps, contradictions, and emerging edge cases to project leads and policy teams - Participate in calibration discussions, challenge interpretations respectfully, and update judgments when stronger reasoning emerges - Maintain accuracy and consistency across long stretches of near-identical evaluations
Requirements - Trained eye from backgrounds such as photography, illustration, design, art direction, generative image work, or visual content moderation - Ability to spot a logo, brand mark, or subtly familiar face in an image - Capacity to hold a rubric steady across long sessions of visually similar images - Willingness to hold a strong opinion without becoming attached to being right - Ability to explain judgment calls clearly enough for audit - Clear separation of personal taste from the customer-supplied standard - Strong written communication - Maturity and sound judgment when handling sensitive imagery and difficult subject matter
Nice to have - A public portfolio, publication credits, or generative work shared on platforms like Civitai, LoRA training, or published prompt work - Experience judging images comparatively through portfolio review, photo competition judging, or creative A/B testing - Formal training in anatomy, color, lighting, or composition - Content moderation or trust and safety experience on image-heavy platforms - Working knowledge of copyright, trademark, and right-of-publicity basics - Prior AI evaluation, RLHF, image labeling, or data annotation work - Familiarity with calibration sessions, inter-rater agreement, or adjudication workflows
Benefits and work setup - Remote within the US - Compensation: $45-55/hour - W-2 employment classification - Schedule: 8 AM to 5 PM Pacific Time, Monday through Friday - Structured exposure limits, content rotation, mandatory reporting protocols, and mental health support for sensitive material