Hiring.Camp

Research Scientist, Agentic Data & Benchmarking

Institute of Foundation Models

·

Jun 8, 2026

Salary
$150k – $450k/yr
Location
Sunnyvale, CA
Workplace
Onsite
Type
Full-time
Department
Research
Experience
2+ years
Education
PhD
Source
Lever

Description

About the Institute of Foundation Models 

The Institute of Foundation Models (IFM) is a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy. 

As part of our team, you'll work at the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You'll help build groundbreaking AI systems with the potential to reshape entire industries, and contribute to establishing MBZUAI as a global hub for high-performance computing and deep learning. 

About the role 

The Agents team trains advanced agentic language models that use reasoning and tool use to complete real tasks on a computer. This is a specialist role at the center of the loop that drives those models: the data we train on and the benchmarks we measure against. 

You'll own the agentic data pipeline end-to-end — sourcing and generating high-quality trajectories, tool-use data, and RL environments — and the evaluation suite that tells us, rigorously and reproducibly, what our agents can actually do. These two halves are inseparable: benchmarks expose where models fail, and targeted data closes the gap. The agents are only as good as the data they learn from and the evals that keep us honest, and this role owns both. 

This is a research scientist position for someone who wants depth in data and measurement rather than breadth across the whole stack. You should be the kind of person who reads through datasets line by line, distrusts a metric until it's been validated, and gets satisfaction from making an eval suite that nobody questions. 

Skills

PythonMachine LearningDeep LearningPyTorch

Similar Jobs

8

AI/ML Research Scientist - Agentic Systems

Peraton · Herndon, VA, US +1 · Onsite

2 months ago

AI/ML Research Scientist - Agentic Systems

Peraton · Herndon, VA, US +2 · Onsite

5 months ago

Research Scientist (Generative & Agentic AI) - Tokyo,Japan.

Appier · Tokyo, Japan +1

3 weeks ago

Research Scientist (Generative & Agentic AI)

Appier · Taipei, Taiwan

1 month ago

Research Scientist/Engineer (Agentic Systems)

White Circle · Paris or London · Hybrid

2 months ago

Senior Agentic AI Research Scientist

Axon · Seattle, Washington, United States +1 · Remote, Hybrid, Onsite

2 months ago

Applied AI Research Scientist - Generative and Agentic AI / Scientifique en IA Appliquée- IA Générative et Agentique

Thales · Quebec City, Canada · Hybrid

1 month ago

Principal Scientist / Associate Director, Agentic AI Research for Materials Science

Lilasciences · Cambridge, MA USA; San Francisco, CA USA

1 month ago