Senior AI Researcher
We are seeking a Senior AI Researcher to lead post-training evaluation, red-teaming, and reinforcement learning (RL) gym audits on open-weight models. The ideal candidate will establish rigorous benchmarking methodologies, evaluate large language models (LLMs) against complex threats like Indirect Prompt Injections (IPI), and construct post-training evaluation pipelines that accurately measure realistic frontier-level security capabilities.
Key Responsibilities
- RL Post-Training & Benchmarking: Execute post-training runs (e.g. GRPO) using mainstream open-weight generalist models against security-focused RL environments, targeting threat vectors like Indirect Prompt Injection (IPI). Reward Diagnostics & Trace Analysis - Analyze live loss curves and rollout traces to identify reward hacking, lazy policy convergence, and flawed or over/under-specified verifiers.
- Task & Environment Auditing: Review tasks and multi-turn environments (including tool use, web navigation, and computer use) for realism, threat model accuracy, data distribution, and dataset balance.
- Performance Reporting (Gym Cards): Generate comprehensive evaluation cards detailing hill-climbing performance uplift across checkpoints, failure modes, tokens/turns per rollout, and task-level success rates.
- Integration & Orchestration: Integrate dockerized environments (e.g., Harbor format) into internal training frameworks, optimizing reset/statefulness semantics, concurrency, and throughput ceilings.
Required Qualifications
- Technical Background: M.S. or Ph.D. in Data Science, Machine Learning, Computer Science, or equivalent practical experience in deep learning.
- RL & Post-Training Expertise: Strong hands-on experience training large-scale models using RL algorithms (e.g. GRPO, PPO) on open-weight architectures.
- AI Security Expertise: Solid understanding of LLM vulnerabilities, red-teaming methodologies, and defensive alignment against IPI attacks.
- Infrastructure Skills: Proficiency in PyTorch, Docker containerization, and distributed training architectures.
- Diagnostic Skills: Ability to analyze agent rollout traces, craft deterministic rubrics/verifiers, and debug complex reward shaping flaws.
Preferred Qualifications
- Prior experience working with standard RL gym formats, such as Harbor.
- Experience evaluating complex agentic workflows in tool-use or web-browser environments.
- Familiarity with evaluating open-weight models similar to Llama or Mistral against adversarial workloads.
Alice is a trust, safety, and security company built for the AI era. We safeguard the communicative technologies people use to create, collaborate, and interact- whether with each other or with machines.
In a world where AI has fundamentally changed the nature of risk, Alice provides end-to-end coverage across the entire AI lifecycle. We support frontier model labs, enterprises, and UGC platforms with a comprehensive suite of solutions: from model hardening evaluations and pre-deployment red-teaming to runtime guardrails and ongoing drift detection.
Alice is widely considered a global leader in online safety and AI security. We have some of the most forward-thinking and passionate minds in the world working to safeguard over 3 billion users across the largest AI and tech platforms.
If you're creative and driven to secure the future of AI, we want to hear from you!
Required Skills
Required Languages
🇬🇧 English