Middle Career Centre | Recruitment
AI Evaluation Scientist - Learning & Cognition
Full Time - Professional Staff Professional Staff Yesterday Requisition ID: 1228 Salary Range: $125,800.00 To $157,300.00 AnnuallyAI Evaluation Scientist, Learning and Cognition
POSITION SUMMARY
We are seeking an AI Evaluation Scientist, Learning and Cognition to design and implement evaluation methods for AI learning agents to be used in the post-secondary education sector.
This Scientist role focuses on establishing criteria for high-quality learning agents and how they should be evaluated. You will translate educational goals into valid constructs, rubrics, test sets, automated evaluators, and release criteria. You will work across conformance testing, LLM-as-judge methods, synthetic learner simulation, adversarial testing, human calibration, and trace analysis of real-world use.
This role requires a combination of investigating how people learn and reason with practical experience in evaluating agentic AI systems. You will be the team's primary source of learning and cognition expertise, working closely with faculty researchers in the field of human-computer interaction and with peers building agents and evaluation harnesses.
KEY RESPONSIBILITIES
Translate educational requirements into valid, testable constructs and rubrics, such as explanation quality, reasoning depth, mastery checking, logical consistency, and resistance to sycophancy;
Distinguish whether an agent conforms to its intended design from whether its behaviour is likely to support learning, and design evaluations for each;
Set evidence-based release criteria for new agents and agent changes;
Design evaluation pipelines for multi-turn agent trajectories, not only single responses;
Build and validate LLM-as-judge evaluators, applying current best practice from agentic evaluation research;
Create libraries of synthetic learner personas grounded in the learning sciences, representing different prior knowledge, common misconceptions, help-seeking behaviours, and interaction patterns;
Run adversarial testing and red teaming to identify inappropriate agent behaviours against safety and learning criteria;
Build regression suites that detect behavioural changes after updates to models, prompts, tools, source material, or agent specifications;
Design and run annotation studies with teachers and students from our user base, including rater recruitment, rubric training, and quality control;
Measure inter-rater reliability and human–machine agreement and decide which judgments each rater group can validly make;
Investigate evaluator failure modes and bias, including model-family bias, position bias, verbosity bias, and evaluator sycophancy;
Communicate results clearly to the engineering team, instructors, partners, and leadership; and,
Other related duties as assigned from time to time.
KEY SUCCESS MEASURES
A reusable evaluation harness is in use across at least two kinds of learning agents;
Synthetic learner and adversarial probe libraries consistently surface meaningful agent failures;
Automated evaluators show demonstrated, documented agreement with teacher and expert ratings;
There is a clear, agreed split between evaluation tasks that are automated and those that need human review;
Regression suites reliably catch behavioural changes before release;
Results are reproducible, auditable, and tied to specific agent and evaluator versions;
Evaluation findings have led to measurable improvements in agent behaviour; and,
Partner pilots are using your agent-behaviour measures in their study designs.
PROFILE OF THE IDEAL CANDIDATE
PhD or Master’s degree with relevant experience in cognitive science, psychology, learning science, human-computer interaction, or computer science/ML with substantial research in human learning or cognition;
2+ years experience working with agentic systems with a focus on contemporary agentic evaluation science;
Familiarity with learning science research on misconceptions, formative assessment, feedback, and productive struggle;
Experience with evaluation and observability tools such as Inspect, Langfuse, or Braintrust;
Experience building and evaluating AI agents, including fluency with agent tracing and application of state-of-the-art ML and LLM tooling;
Proficient in Python for research work, including building evaluation pipelines, analyzing data, and working with LLM APIs and evaluation tooling;
Experience with statistical, LLM-as-judge methods, agent trajectory evaluation, adversarial testing, red teaming, privacy and safety evaluations, and the creation of synthetic data sets to support evaluation of agentic systems;
Experience designing evaluation datasets, benchmarks, annotation studies, or controlled experiments;
Strong understanding of quantitative and qualitative evaluation methods, experimental design, statistical comparison, and inter-rater reliability;
Ability to translate qualitative behavioural requirements into precise, measurable evaluation criteria;
Strong verbal and written communication skills and the ability to work effectively with ML specialists, software developers, scientists, instructors, and subject-matter experts; and,
A passion for learning and continuous growth.
TOTAL REWARDS: The expected salary for this position will be $125,800 - $157,300 per year, plus benefits if applicable. The final salary offer will reflect the successful candidate's experience, skills, and qualifications, in alignment with the Vector Institute's Compensation Policy and may differ from above.
The Vector Institute’s Total Rewards approach extends beyond traditional compensation and benefits. Full-time employees are eligible for a comprehensive suite of supports that recognize and value employees, including vacation time, floater days, GRRSP, a Health Spending Account, a Summer Hours program, and flexible work arrangements.
POSITION STATUS: This posting is for a new vacancy.
USE OF ARTIFICIAL INTELLIGENCE: Vector may use both internal and external third party AI-based tools to assist in the screening of applications for this posting. Any data collected will be used solely for recruitment purposes and handled in accordance with Vector’s External Privacy Policy and Use of AI-Based Tools in Recruitment and Selection Policy.
INCLUSION AND EQUAL OPPORTUNITY EMPLOYMENT: Vector believes AI powers possibility by advancing cutting-edge research and translating it into real-world impact through collaboration with research, industry, and government. Vector is committed to fostering a diverse and inclusive culture that reflects its values.
The Vector Institute welcomes applications from all qualified candidates, including those who are Indigenous, 2SLGBTQIA+, racialized persons/visible minorities, women, and people with disabilities.
JOIN OUR COMMUNITY: Check out the Vector Institute’s Careers Page to explore open opportunities at Vector and Follow Vector on X, LinkedIn, and Bluesky to stay connected with the latest developments in Ontario's AI ecosystem and the Vector Institute.
Required Skills
Required Languages
🇬🇧 English