Hiring.Camp

AI Evaluation Engineer (QA)

Appnovation

·

Today

Location
Toronto · São Paulo, Brazil · Toronto, Canada
Type
Full-time
Department
Global Services Technology
Source
Greenhouse

Description

About us

Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth.  Bold ambition. Practical action. Endless possibilities.

Technology is foundational to all of Appnovation’s offerings, from consulting to digital innovation, to digital product and service creation. The technology department is focused on delivering software solutions that enable rich consumer experiences, from mobile and web applications to advanced analytics and machine learning, to content and engagement management service enablement platforms.

Inherent throughout our tech capabilities is deep expertise in the Software Development Life Cycle, a drive for creativity, a passion for the craft, and collaboration with other disciplines - all foundational ingredients in successful digital experiences and client partnerships.

As a QA / AI Evaluation Engineer, you will join a highly motivated and experienced team in a forward-leaning role that proves the platform actually improves answer quality. You will run evaluations at scale — from small human-UAT batches up to millions of automated evals — statistically measure factual grounding and accuracy lift, and build the metrics framework that shows how much better our answers get over time. We are looking for people who can bring a strong, solution-focused mindset and contribute to quality standards, best practices and get things done.

KEY RESPONSIBILITIES

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., “answer is X% supported by source content / Y% better”).
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

QUALIFICATIONS

  • Bachelor’s Degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.

WHO YOU ARE

  • You think about how to scale, automate and operate, not just how to build a solution to an immediate problem
  • You understand lean thinking
  • You set high standards for code quality, performance/scalability and security and seek continuous improvement
  • You have solid analytical, problem solving and decision-making skills
  • You have customer first mindset and a devotion to customer service
  • You engage and build positive internal and external client relationships, while managing multiple initiatives, often with competing priorities
  • You have strong self-initiative, passion, interpersonal, oral and written communication and collaboration skills with the ability to work, influence and make an impact in a cross-functional environment with all levels of the organization
  • You are responsive and thrive in a fast-paced diverse high-performance environment with rapidly changing business needs
  • You actively seek out things outside your comfort zone with the ability to rapidly learn and take advantage of new concepts, business models, and technologies
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred
Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.
 
At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital.

Accommodations are available upon request throughout the recruitment process.

Skills

PythonCI/CDMachine LearningCustomer Service

Similar Jobs

30

AI Evaluation Engineer

Zafin · Toronto

1 month ago

AI Evaluation Engineer

Judi Health · Charlotte, North Carolina, United States; Denver, Colorado, United States; New York, New York, United States · Remote

1 month ago

AI Engineer (Evaluation)

NTU Singapore · NTU Main Campus, Singapore · onsite

2 months ago

AI Engineer, Evaluation

Distyl · San Francisco · Hybrid

2 months ago

AI Evaluation Engineer

Yes Energy · Bucharest, Romania +1 · Hybrid

2 months ago

AI Evaluation Engineer (QA)

Appnovation · San José, Medellin, Bogota, Mexico City, Buenos Aires +2

Today

AI Evaluation Engineer (QA)

Appnovation · London +2

Today

AI Evaluation Engineer (QA)

Appnovation · Montreal +2

Today

AI Evaluation Engineer (QA)

Appnovation · New York, Austin, Miami, Dallas, City of Phoenix +2

Today

AI Evaluation Engineer (QA)

Appnovation · São Paulo +2

Today

Senior Rust Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)

G2I · Miami +35 · Remote

1 week ago

Senior Python Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)

G2I · Miami +35 · Remote

1 week ago

Senior Go Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)

G2I · Miami +35 · Remote

1 week ago

Software Engineer, AI Evaluation

Nuna · San Francisco · Hybrid

1 week ago

Senior Applied AI Evaluation Engineer - RAN Ticket Intelligence

Parallelwireless · Kfar Saba · Onsite

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · São Paulo, SP, Brazil · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Canberra, ACT, Australia · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Madrid, MD, Spain · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Warsaw, Masovian Voivodeship, Poland · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Rome, Lazio, Italy · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Paris, IDF, France · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Ottawa, ON, Canada · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · London, England, United Kingdom · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Berlin, BE, Germany · Remote

3 weeks ago

Copy of CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Hyderabad, TS, India · Remote

3 weeks ago

CFD & Aerodynamic Engineer (AI Training & Evaluation)

Lifted · Houston, TX, United States · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · London, England, United Kingdom · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Ottawa, ON, Canada · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Warsaw, Masovian Voivodeship, Poland · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Jakarta, Jakarta, Indonesia · Remote

3 weeks ago
AI Evaluation Engineer (QA) at Appnovation | Hiring.Camp