Remote job
AI Red Teamer (LLM Generalist) - Remote
Job details
About this role
Role overview This is a generalist adversarial testing role focused on stress-testing large language models. Rather than checking correctness, the work involves designing creative prompts that expose unsafe content, bias, broken guardrails, hallucinations, prompt injection weaknesses, and other unexpected behaviors. Testing spans content safety, CBRN, cybersecurity, persuasion and influence operations, child safety, self-harm, over-companiership, and regulatory compliance, across text, image, voice, and agentic capabilities.
Responsibilities - Craft creative prompts and multi-turn scenarios to probe AI guardrails across diverse risk categories. - Discover ways around safety filters, restrictions, and defenses using jailbreak, evasion, and prompt injection techniques. - Evaluate and score model responses against structured harm taxonomies and severity rubrics. - Document experiments clearly, including hypotheses, methods tried, and what they revealed. - Review and refine adversarial prompts produced by other team members and contribute to harm taxonomy, calibration, and inter-rater reliability work. - Collaborate with engineers, data scientists, and researchers to share findings and strengthen defenses, while staying current on jailbreaks and emerging model behaviors.
Requirements - Strong hands-on experience using multiple LLMs such as ChatGPT, Claude, Gemini, and open-source models. - Intuition for crafting adversarial prompts, with familiarity in jailbreak or evasion techniques as a strong plus. - Creative, adversarial problem-solving skills paired with clear and thoughtful written communication. - Strong ethical judgment and the ability to separate adversarial thinking from personal values. - Self-directed and collaborative, with comfort operating in feedback-heavy environments and with frequent experimental failure. - Ability to work professionally and sustainably with regularly disturbing content as part of structured testing.
Nice to have - Familiarity with Python or other scripting languages. - Experience with LLM APIs or evaluation tooling, structured data annotation, or rubric-based scoring. - Prior work in trust and safety, content moderation, QA, or security research. - Subject matter expertise in high-risk domains such as cybersecurity, chemistry, biology, medicine, law, or finance. - Creative background in writing, visual art, improv, puzzle design, or similar fields, with a tendency to dive deep into unusual interests.
Benefits and work setup - Remote position based in the USA, structured as a 40-hour-per-week contract engagement. - Support resources are available to help candidates engage sustainably with harmful content during structured adversarial testing.