Back to jobs

AI SAFETY SPECIALIST - FULLY REMOTE | UPTO $25/HR

mercor
Part-timemid€17-25/hour

Job description

About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Malay Type: Contract Compensation: $17–$25/hour Location: Remote Role Responsibilities • Red team conversational AI models and agents to identify jailbreaks, prompt injections, and misuse cases. • Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks. • Apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing. • Document reproducibly to produce reports, datasets, and attack cases that customers can act on. • Review AI outputs on sensitive topics such as bias and misinformation, with optional participation in higher-sensitivity projects. Qualifications Must-Have • Fluent Language Skills Required: English & Malay • Prior red teaming experience in AI adversarial work , cybersecurity, or socio-technical probing. • Ability to explain risks clearly to technical and non-technical stakeholders. Preferred • Experience in Adversarial ML : jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction. • Background in Cybersecurity : penetration testing, exploit development, reverse engineering. • Expertise in socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing. • Creative probing skills: psychology, acting, writing for unconventional adversarial thinking. Application Process (Takes 20–30 mins to complete) • Upload resume • AI interview based on your resume • Submit form Resources & Support • For details about the interview process and platform information, please check: https://talent.docs.mercor.com/welcome • For any help or support, reach out to: support@mercor.com PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Skills

red teamingAI adversarial workcybersecuritysocio-technical probingAdversarial MLjailbreak datasetsprompt injectionRLHFDPO attacksmodel extractionpenetration testingexploit developmentreverse engineeringharassment/disinfo probingabuse analysisconversational AI testingpsychologyactingwriting