Hiring.Camp

AI Evaluation Engineer

Dialpad

·

2 days ago

Location
Kitchener, Canada · Vancouver, British Columbia, Canada
Department
214 - AI Engineering
Experience
3+ years
Education
Master
Source
Greenhouse

Description

About Dialpad
Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.

Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.

Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile.

Being a Dialer
At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more.

We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves.

We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic.

Your role
As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions.

This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.

What you’ll do 

  • You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
  • You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
  • You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
  • You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.
  • You will develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams.
  • You will investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis.
  • You will collaborate with cross-functional teams, including applied science, engineering, and Product QA.

Skills you’ll bring 

  • Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
  • 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
  • Experience designing structured test strategies across manual and automated workflows.
  • Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
  • Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
  • Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
  • Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.

For exceptional talent based in Ontario, Canada the target base salary range for this position is posted below. Our salary ranges are determined by role, level, and location. The range displayed on each job posting reflects the target range for new hire salaries for the position. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your preferred location during the hiring process. Please note that the compensation details listed in Ontario role postings reflect the base salary only, and do not include bonus, equity, or benefits.

Ontario Salary Range
$104,000$119,750 CAD

Why Join Dialpad

  • Work at the center of the AI transformation in business communications
  • Build and ship agentic AI products that are redefining how companies operate
  • Join a team where AI amplifies every employee’s impact
  • Competitive salary, comprehensive benefits, and real opportunities for growth

We believe in investing in our people. Dialpad offers competitive benefits and perks, cutting-edge AI tools, and a robust training program that help you reach your full potential. We have designed our offices to be inclusive, offering a vibrant environment to cultivate collaboration and connection. Our exceptional culture, repeatedly recognized as a Great Place to Work, ensures that every employee feels valued and empowered to contribute to our collective success.

Don’t meet every single requirement? If you’re excited about this role and possess the fundamental traits, drive, and strong ambition we seek, but your experience doesn’t meet every qualification, we encourage you to apply. 

 Dialpad is an equal-opportunity employer. We are dedicated to creating a community of inclusion and an environment free from discrimination or harassment.

Skills

NLP

Similar Jobs

30

AI Evaluation Engineer

Zafin · Toronto

2 days ago

AI Evaluation Engineer

GovWorx · United States · Remote

1 week ago

AI Evaluation Engineer

Judi Health · Charlotte, North Carolina, United States; Denver, Colorado, United States; New York, New York, United States · Remote

1 week ago

AI Evaluation Engineer

FNZ is committed to opening · IN Gurugram, India

3 weeks ago

AI Engineer (Evaluation)

NTU Singapore · NTU Main Campus, Singapore · onsite

1 month ago

AI Engineer, Evaluation

Distyl · San Francisco · Hybrid

1 month ago

AI Evaluation Engineer

Yes Energy · Bucharest, Romania +1 · Hybrid

1 month ago

AI Evaluation Engineer

FNZ is committed to opening · IN Pune, India

1 month ago

AI Engineer (Evaluation)

NTU Singapore · NTU Main Campus, Singapore

3 months ago

AI Evaluation Engineer

Distyl · San Francisco +1 · Hybrid

3 months ago

Staff Software Development Test Engineer - AI Evaluation

Tekion · Bengaluru, Karnataka, India

2 weeks ago

AI Quality & Evaluation Engineer, AI product

Woven by Toyota · Tokyo · Hybrid

1 month ago

Senior Software Engineer — AI Evaluation & Benchmarks (Python)

G2I · Miami +206 · Remote

2 months ago

AI Engineer, Evaluation & Quality - 11319

Coupa · Bangalore, India · Remote

4 months ago

Applied AI, Evaluation Engineer

Mistral · Paris · Onsite

6 months ago

Staff ML Engineer - Embodied AI Evaluation Foundations

General Motors · GM Automation - Sunnyvale - GM Automation - Sunnyvale, United States of America · Remote

4 days ago

[AI Research Div.] Foundation Model Evaluation Engineer - 독자 AI 파운데이션 모델 (2년 이상 / 인턴)

Krafton · Seoul

1 week ago

Senior AI Engineer - Agentic AI Evaluation

Resaro AI · Munich, Bavaria, Germany

2 weeks ago

Senior AI Engineer - LLM Evaluation

AutoDesk · APAC - India - Bengaluru - Sunriver

1 month ago

Senior Applied AI Engineer – Prompting & Evaluation

Kerridge Commercial Systems · Johannesburg, SA +1

1 month ago

Applied AI Engineer – Prompting & Evaluation

Kerridge Commercial Systems · Johannesburg, SA +1

1 month ago

Applied AI Engineer – Prompting & Evaluation

Kerridge Commercial Systems · Tyne and Wear, UK +9

1 month ago

Senior Applied AI Engineer – Prompting & Evaluation

Kerridge Commercial Systems · Tyne and Wear, UK +10

1 month ago

Software Engineer, AI Data & Evaluation

Mercor · San Francisco · Onsite

1 month ago

Senior AI Engineer, Agentic Evaluation & V&V

Slingshotaerospace · Remote · Remote

2 months ago

Senior ML/AI Software Engineer – Evaluation Insights

General Motors · GM Automation - Sunnyvale - GM Automation - Sunnyvale, United States of America · Remote

5 months ago

AI Engineer, Agents & Evaluation

Guild.ai · San Francisco · Hybrid

7 months ago

Senior/Staff Machine Learning Engineer - Health Evaluation - AI Teams (x/f/m)

Doctolib · Paris, Paris, France · Remote

8 months ago

Copy of Senior Structural Engineer (OpenSees | AI Training & Evaluation)

Lifted · Lisbon, Lisbon, Portugal · Remote

3 days ago

Copy of Senior Structural Engineer (OpenSees | AI Training & Evaluation)

Lifted · Santiago, Santiago Metropolitan Region, Chile · Remote

3 days ago