RVC
JobsFor Employers
  1. Jobs
  2. /
  3. Middle Career Centre | Recruitment

Middle Career Centre | Recruitment

Vector Institute | Artificial intelligence research
1h 55m ago
On-site
Full Time
IC
125.8k - 157.3k CAD/yr≈ $88.4k - $110.5k USD/yr

AI Evaluation Scientist - Learning & Cognition

Full Time - Professional Staff Professional Staff Yesterday Requisition ID: 1228 Salary Range: $125,800.00 To $157,300.00 Annually

AI Evaluation Scientist, Learning and Cognition

POSITION SUMMARY

We are seeking an AI Evaluation Scientist, Learning and Cognition to design and implement evaluation methods for AI learning agents to be used in the post-secondary education sector.

This Scientist role focuses on establishing criteria for high-quality learning agents and how they should be evaluated. You will translate educational goals into valid constructs, rubrics, test sets, automated evaluators, and release criteria. You will work across conformance testing, LLM-as-judge methods, synthetic learner simulation, adversarial testing, human calibration, and trace analysis of real-world use.

This role requires a combination of investigating how people learn and reason with practical experience in evaluating agentic AI systems. You will be the team's primary source of learning and cognition expertise, working closely with faculty researchers in the field of human-computer interaction and with peers building agents and evaluation harnesses.

KEY RESPONSIBILITIES

  • Translate educational requirements into valid, testable constructs and rubrics, such as explanation quality, reasoning depth, mastery checking, logical consistency, and resistance to sycophancy;

  • Distinguish whether an agent conforms to its intended design from whether its behaviour is likely to support learning, and design evaluations for each;

  • Set evidence-based release criteria for new agents and agent changes;

  • Design evaluation pipelines for multi-turn agent trajectories, not only single responses;

  • Build and validate LLM-as-judge evaluators, applying current best practice from agentic evaluation research;

  • Create libraries of synthetic learner personas grounded in the learning sciences, representing different prior knowledge, common misconceptions, help-seeking behaviours, and interaction patterns;

  • Run adversarial testing and red teaming to identify inappropriate agent behaviours against safety and learning criteria; 

  • Build regression suites that detect behavioural changes after updates to models, prompts, tools, source material, or agent specifications;

  • Design and run annotation studies with teachers and students from our user base, including rater recruitment, rubric training, and quality control;

  • Measure inter-rater reliability and human–machine agreement and decide which judgments each rater group can validly make;

  • Investigate evaluator failure modes and bias, including model-family bias, position bias, verbosity bias, and evaluator sycophancy;

  • Communicate results clearly to the engineering team, instructors, partners, and leadership; and,

  • Other related duties as assigned from time to time.

KEY SUCCESS MEASURES

  • A reusable evaluation harness is in use across at least two kinds of learning agents;

  • Synthetic learner and adversarial probe libraries consistently surface meaningful agent failures;

  • Automated evaluators show demonstrated, documented agreement with teacher and expert ratings;

  • There is a clear, agreed split between evaluation tasks that are automated and those that need human review;

  • Regression suites reliably catch behavioural changes before release;

  • Results are reproducible, auditable, and tied to specific agent and evaluator versions;

  • Evaluation findings have led to measurable improvements in agent behaviour; and,

  • Partner pilots are using your agent-behaviour measures in their study designs.

PROFILE OF THE IDEAL CANDIDATE

  • PhD or Master’s degree with relevant experience in cognitive science, psychology, learning science, human-computer interaction, or computer science/ML with substantial research in human learning or cognition;

  • 2+ years experience working with agentic systems with a focus on contemporary agentic evaluation science; 

  • Familiarity with learning science research on misconceptions, formative assessment, feedback, and productive struggle;

  • Experience with evaluation and observability tools such as Inspect, Langfuse, or Braintrust; 

  • Experience building and evaluating AI agents, including fluency with agent tracing and application of state-of-the-art ML and LLM tooling;

  • Proficient in Python for research work, including building evaluation pipelines, analyzing data, and working with LLM APIs and evaluation tooling;

  • Experience with statistical, LLM-as-judge methods, agent trajectory evaluation, adversarial testing, red teaming, privacy and safety evaluations, and the creation of synthetic data sets to support evaluation of agentic systems;

  • Experience designing evaluation datasets, benchmarks, annotation studies, or controlled experiments;

  • Strong understanding of quantitative and qualitative evaluation methods, experimental design, statistical comparison, and inter-rater reliability;

  • Ability to translate qualitative behavioural requirements into precise, measurable evaluation criteria;

  • Strong verbal and written communication skills and the ability to work effectively with ML specialists, software developers, scientists, instructors, and subject-matter experts; and,

  • A passion for learning and continuous growth.

TOTAL REWARDS: The expected salary for this position will be $125,800 - $157,300 per year, plus benefits if applicable. The final salary offer will reflect the successful candidate's experience, skills, and qualifications, in alignment with the Vector Institute's Compensation Policy and may differ from above.

The Vector Institute’s Total Rewards approach extends beyond traditional compensation and benefits. Full-time employees are eligible for a comprehensive suite of supports that recognize and value employees, including vacation time, floater days, GRRSP, a Health Spending Account, a Summer Hours program, and flexible work arrangements. 

POSITION STATUS: This posting is for a new vacancy.

USE OF ARTIFICIAL INTELLIGENCE: Vector may use both internal and external third party AI-based tools to assist in the screening of applications for this posting. Any data collected will be used solely for recruitment purposes and handled in accordance with Vector’s External Privacy Policy and Use of AI-Based Tools in Recruitment and Selection Policy.  

INCLUSION AND EQUAL OPPORTUNITY EMPLOYMENT: Vector believes AI powers possibility by advancing cutting-edge research and translating it into real-world impact through collaboration with research, industry, and government. Vector is committed to fostering a diverse and inclusive culture that reflects its values.

The Vector Institute welcomes applications from all qualified candidates, including those who are Indigenous, 2SLGBTQIA+, racialized persons/visible minorities, women, and people with disabilities.

JOIN OUR COMMUNITY: Check out the Vector Institute’s Careers Page to explore open opportunities at Vector and Follow Vector on X, LinkedIn, and Bluesky to stay connected with the latest developments in Ontario's AI ecosystem and the Vector Institute.

Required Skills

pythonlangfuse

Required Languages

🇬🇧 English

Related searches

  • Data Science Jobs
  • Mid-Level Jobs
1 jobs
Sort by
1h 55m ago

Middle Career Centre | Recruitment

Vector Institute·Artificial intelligence research
125.8k - 157.3k CAD/yr≈ $88.4k - $110.5k USD/yr
pythonlangfuse
🏢On-site
Data Scientist (AI/ML)
No similar jobs match these filters. Try changing or clearing a filter.
Popular job searches ▸
Roles
Python JobsJava JobsReact JobsTypeScript JobsGo JobsDevOps JobsNode.js JobsRemote C# and .NET JobsData Science JobsProduct Manager JobsDesign Jobs
Countries
United StatesUnited KingdomCanadaUS & EMEAGermanyPolandSpainNetherlandsPortugal
Experience
Senior JobsMid-Level JobsJunior Jobs
R© 2026 RVC Globalbuild 5ac55c1a
AboutJobs in Claude & ChatGPTPricingContactPrivacyCookiesRefundsTerms & ConditionsFor Employers