AI SAFETY SPECIALIST - FULLY REMOTE | UPTO $45/HR
mercor
Part-timemid€29-45/hour
Job description
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: AI Safety Experts — English & Portuguese (global)
Type: Contract
Compensation: $29–$45/hour
Location: Remote
Role Responsibilities
• Red team conversational AI models and agents by performing jailbreaks, prompt injections, misuse cases, and bias exploitation.
• Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
• Apply structure by following taxonomies, benchmarks, and playbooks to maintain consistent testing.
• Document reproducibly by producing reports, datasets, and attack cases that customers can act on.
• Work independently and asynchronously to improve AI model performance and ensure safety.
Qualifications
Must-Have
• Native fluency in English and Portuguese (global, excluding Brazilian Portuguese).
• Prior red teaming experience in AI adversarial work , cybersecurity , or socio-technical probing .
• Strong communication skills to explain risks clearly to technical and non-technical stakeholders.
Preferred
• Experience in Adversarial ML : jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
• Cybersecurity skills: penetration testing, exploit development, reverse engineering.
• Socio-technical risk expertise: harassment/disinfo probing, abuse analysis, conversational AI testing .
• Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.
Application Process (Takes 20–30 mins to complete)
• Upload resume
• AI interview based on your resume
• Submit form
Resources & Support
• For details about the interview process and platform information, please check: https://talent.docs.mercor.com/welcome
• For any help or support, reach out to: support@mercor.com
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
Skills
Red teamingAI adversarial workCybersecuritySocio-technical probingAdversarial MLJailbreak datasetsPrompt injectionRLHF/DPO attacksModel extractionPenetration testingExploit developmentReverse engineeringHarassment/disinfo probingAbuse analysisConversational AI testingPsychologyActingWriting