Remote job
AI Red Teamer (Remote)
Job details
About this role
Role overview
This position focuses on adversarial testing of large language models, using creative prompt design to expose safety vulnerabilities before they reach end users. You will probe guardrails across text, voice, and agentic capabilities, working across risk categories such as content safety, CBRN, cybersecurity, persuasion, child safety, and regulatory compliance. The work directly supports AI safety research and helps strengthen frontier models against misuse.
Responsibilities
- Design adversarial prompts and multi-turn scenarios that stress test models across diverse risk categories - Discover methods for bypassing safety filters using jailbreak, evasion, and prompt injection techniques - Evaluate and score model responses against structured harm taxonomies and severity rubrics - Document experimental findings, including attempts, reasoning, and resulting model behavior - Review and refine adversarial prompts contributed by other team members - Collaborate with engineers, data scientists, and researchers to share findings and strengthen defenses
Requirements
- Hands-on experience working with multiple large language models, including commercial and open-source systems - Intuition for crafting adversarial prompts, with familiarity in jailbreak or evasion techniques - Creative, adversarial problem-solving skills paired with strong ethical judgment - Clear and thoughtful written communication - Self-directed and collaborative, with comfort working in feedback-heavy environments - Willingness to engage regularly with potentially disturbing content as part of structured testing
Nice to have
- Familiarity with Python or other scripting languages - Experience using LLM APIs or evaluation tooling - Background in trust and safety, content moderation, QA, or security research - Subject-matter expertise in a high-risk domain such as cybersecurity, chemistry, biology, medicine, law, or finance
Benefits and work setup
- Remote position within the United States; California residents are not eligible for this role - Full-time W-2 hourly employment at 40 hours per week - This role involves regular exposure to harmful content as part of structured adversarial testing, with support resources available