Job – AI Research Engineer – nineDots
AI Research Engineer
Posted on Oct 6, 2026 in Software Engineering (AI Engineering)
Dublin
3 months left
EUR 100000 - 140000 / Year
EUR 100000 - 120000 / Year
Dublin city centre location
5 days onsite
Small, ambitious research team
Work on frontier models and benchmarking
Specialist AI research role focused on benchmarking, evaluation and adversarial testing of frontier LLMs within a fast-moving Dublin research team working fully onsite.
Overview
This is a deliberately specialised role for someone deeply focused on frontier AI models rather than product integrations.
You will join a small research team in Dublin working on training and evaluation environments for leading AI labs. The team moves quickly and works hard, so this role will suit someone who actively wants a start-up environment rather than a traditional 9-to-5.
This is a 5 days per week onsite role in Dublin city centre.
What You'll Be Doing
- Designing and building benchmarks for frontier LLMs
- Comparing how different models perform and behave
- Building difficult tasks and datasets to expose model strengths and weaknesses
- Adversarially testing models and finding where they break
- Designing rigorous evaluation methodologies
- Working with synthetic training and evaluation data
- Investigating model reasoning, reliability and unexpected behaviour
- Working on post-training and reinforcement learning
- Running experiments across multiple frontier models
What You'll Need to Succeed
Experience in one or more of the following:
- Building or contributing to recognised or public LLM benchmarks
- LLM benchmarking or model capability evaluation
- Adversarial LLM testing or red teaming
- AI safety or model safety research
- LLM post-training or reinforcement learning
- Synthetic data generation for model training or evaluation
- Designing evaluation frameworks, metrics or model judges
- Research into LLM behaviour and failure modes
- Working deeply across multiple models such as Claude, GPT, Gemini, Llama or DeepSeek
Strong research credentials are highly valued. This could include:
- A PhD in a relevant area
- Significant research experience
- Publications at conferences such as NeurIPS, ICML or ICLR
- Demonstrably strong work building benchmarks and evaluation systems
Most importantly, this role is looking for people who have actually done this work, rather than simply used the terminology.
Likely Not a Fit If Your Experience Is Mainly
- RAG
- LangChain
- Vector databases
- Prompt engineering
- Building chatbots
- Agent orchestration
- Calling LLM APIs
- Adding GenAI features to existing products
Without deeper model evaluation, benchmarking or research experience, this role is unlikely to be a fit.
What's In It for You
- Salary of EUR 100000 - 120000 / Year
- Opportunity to work on frontier model benchmarking and evaluation
- Small, ambitious team environment
- Exposure to multiple leading frontier models
- Central Dublin location with fully onsite collaboration
If you are the kind of person who sees a new frontier model released and immediately wants to test it, break it, compare it and understand why it behaves differently, this role could be a strong match.
Required Skills
Required Languages
🇬🇧 English