Hiring.Camp

Research Engineer – Benchmarking, Evals & Failure Analysis

RFS Group

·

Apr 21, 2026

Salary
$130k – $400k
Workplace
Onsite
Type
Full-time
Department
Engineering
Source
RecruiterFlow

Description

Research Engineer – Benchmarking, Evals & Failure Analysis

Location: San Francisco
Company Stage: Late-Stage / Series C (AI / Applied ML)
Office Type: Onsite (5 Days a Week)
Salary: $130,000 – $400,000 + Equity


Company Description

This fast-growing AI company is operating at the forefront of applied machine learning and labor transformation. By partnering with leading AI labs and enterprises, they are building systems that combine human expertise with cutting-edge AI to improve model performance and unlock new categories of work. With strong revenue, scale, and backing from top-tier investors, they are shaping how frontier models are trained, evaluated, and deployed in real-world environments.


What You Will Do
  • Design and implement benchmarking systems to evaluate model capabilities such as tool use, reasoning, and agent behavior
  • Build and operate end-to-end evaluation pipelines, including scoring systems, dashboards, and reporting infrastructure
  • Conduct systematic failure analysis on model outputs, identifying key failure modes and translating them into actionable improvements
  • Develop rubrics, evaluators, and scoring frameworks that balance rigor with scalability (human + automated evaluation)
  • Partner with research and applied AI teams to align evaluation systems with training and product goals
  • Analyze data quality and performance trends to inform model training, data generation, and post-training strategies
  • Own evaluation and benchmarking systems in a fast-paced, high-iteration environment

Ideal Background
  • Strong applied AI or ML engineering experience, particularly in model evaluation, benchmarking, or failure analysis
  • Hands-on experience building or running LLM evaluation systems, benchmarks, or experimentation pipelines
  • Strong coding ability (Python or similar) with experience building production-quality systems
  • Solid understanding of data structures, algorithms, and backend systems
  • Experience working with APIs, databases (SQL/NoSQL), and cloud infrastructure
  • Ability to reason deeply about model behavior, experimental results, and system performance
  • Comfortable operating in ambiguous, high-ownership environments with rapid iteration cycles

Preferred
  • Experience working on post-training, RL, or evaluation teams at AI labs or AI-first companies
  • Familiarity with LLM evaluation techniques, benchmarking frameworks, or agent evaluation systems
  • Experience with synthetic data generation, rubric design, or reward modeling workflows
  • Publications or research experience in ML, especially in evaluation or benchmarking
  • Exposure to large-scale experimentation systems or model performance tracking infrastructure

Compensation and Benefits
  • Competitive base salary ($130K – $500K) + meaningful equity
  • Relocation and housing support available
  • Monthly meal stipend and premium wellness perks (e.g., fitness membership)
  • Comprehensive health insurance
  • Opportunity to work directly with frontier AI labs and influence model development at the cutting edge

This is a high-impact role at the intersection of engineering and applied AI research, ideal for candidates excited about defining how next-generation models are evaluated, improved, and deployed at scale.

 

Skills

PythonSQLMachine Learning

Similar Jobs

30

Research Engineer

Quantiphi · IN KA Bengaluru, India +1 · Hybrid

2 days ago

Research Engineer

Psu · Penn State University Park, United States of America · Remote, Hybrid

6 days ago

Research Engineer

Neo4J · London

1 week ago

Research Engineer

Greptile · San Francisco · Onsite

1 week ago

Research Engineer

CNI Careers · AL Fort Novosel G Farrel Rd 6901, United States of America · Onsite

1 week ago

Research Engineer

Tessera Labs · San Jose Office (HQ) +1

1 week ago

Research Engineer

E Ink · Billerica, MA

2 weeks ago

Research Engineer

Ford · Dearborn, MI,US, US

3 weeks ago

Research Engineer

USA01 - USA Automotive · Dearborn, MI, United States, US · Hybrid

3 weeks ago

Research Engineer

USC is · Los Angeles, CA - Health Sciences Campus, United States of America

3 weeks ago

Research Engineer

Firecrawl · San Francisco, California, US · Hybrid, Onsite

4 weeks ago

Research Engineer

Verisign · Reston,Virginia,United States

4 weeks ago

Research Engineer

The LYCRA Company · WAYNESBORO, VA

4 weeks ago

Research Engineer

Antimetal · HQ - NYC · Onsite

1 month ago

Research Engineer

Imc · Sydney, Australia

1 month ago

Research Engineer

Imc · Hong Kong, Hong Kong

1 month ago

Research Engineer

Decagon · San Francisco · Onsite

1 month ago

Research Engineer

Superannotate · San Francisco · Hybrid

1 month ago

Research Engineer

Awetomaton · Beavercreek, OH +1

1 month ago

Research Engineer

Console · San Francisco (On-site) · Onsite

1 month ago

Research Engineer

Helm.ai · Remote - US · Remote

2 months ago

Research Engineer

Helm.ai · Remote - Canada · Remote

2 months ago

Research Engineer

Kog AI · Paris, France · Hybrid

2 months ago

Research Engineer

Datologyai · Redwood City · Remote, Onsite

3 months ago

Research Engineer

NowSecure · Remote

3 months ago

Research Engineer

AutoDesk · AMER - United States - Massachusetts - Boston - Drydock, United States of America · Hybrid

3 months ago

Research Engineer

SID · San Francisco, California, US

3 months ago

Research Programmer

Upenn · Smilow Center for Translational, United States of America

3 months ago

Research Programmer

Upenn · Smilow Center for Translational, United States of America

3 months ago

Research Engineer

Thomson Reuters · United States of America, Eagan, Minnesota +6 · Hybrid

4 months ago