Hiring.Camp

Sr. Evaluation Engineer

LogicMonitor

Location
San Francisco, CA
Workplace
Hybrid
Department
IT
Experience
5+ years
Visa
Not sponsored

Description

About Us:  

We love going to work and think you should too. Our team is dedicated to trust, customer obsession, agility, and striving to be better everyday. These values serve as the foundation of our culture, guiding our actions and driving us towards excellence. We foster a culture of performance and recognition, allowing us to transform growth as we enable our employees to do the best work of their careers.

This role is open to candidates based in or near San Francisco, CA. At LogicMonitor, we hire within our Centers of Energy—vibrant locations where our teams connect, collaborate, and innovate.

To learn more about life at LogicMonitor, check out our Careers Page.

What You'll Do:

LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise.

Our customers love LogicMonitor's ability to bring cloud and traditional IT together into one view, as seen in minimal churn rates, expansion business, and exciting new customer references. In fact, LogicMonitor has received the highest Net Promoter Score of any IT Infrastructure Management provider. LogicMonitor also boasts high employee satisfaction. We have been certified as a Great Place To Work®, and named one of BuiltIn's Best Places to Work for the seventh year in a row! 

Edwin AI is LogicMonitor’s AI-powered observability and incident intelligence platform. It helps enterprise operations teams investigate incidents, identify root causes, recommend remediation, and automate operational workflows.

As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for AI agents, retrieval systems, tool integrations, and complex investigation workflows.

Here's a closer look at this key role:

  • Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness.
  • Build offline and online evaluation pipelines in Python and integrate them with CI/CD, experimentation, model selection, prompt iteration, and release gating.
  • Lead the creation and maintenance of golden datasets and regression suites using alerts, events, metrics, logs, traces, topology, configuration data, incident timelines, change records, ITSM workflows, and historical investigation outcomes.
  • Build representative, customer-specific scenarios covering different technologies, failure modes, operational patterns, and environmental constraints.
  • Use human-authored and AI-assisted methods to generate regression, edge, adversarial, rare, ambiguous, and incomplete-context test cases.
  • Treat evaluation datasets and test suites as first-class components that evolve alongside Edwin AI.
  • Design step-level and trajectory-level evaluations for multi-step and multi-agent workflows.
  • Assess both final outcomes and intermediate behavior, including planning, reasoning consistency, retrieval, evidence use, tool selection, tool parameters, state transitions, escalation decisions, and human-in-the-loop approvals.
  • Identify whether failures originate from models, prompts, retrieval, data quality, tools, agent logic, orchestration, or infrastructure.
  • Evaluate capabilities including incident investigation, on-call assistance, operational question answering, change-impact analysis, remediation recommendations, automated resolution, infrastructure operations, and ITSM and observability integrations.
  • Design and calibrate LLM-based graders against expert human judgment.
  • Monitor AI quality and behavioral drift in production, and convert failures and customer feedback into new tests and safeguards.
  • Establish evaluation-driven development practices and mentor other engineers.
What You'll Need:
  • 5+ years of experience in software engineering, machine learning, applied AI, or a related field.
  • Strong Python engineering skills and experience building production systems.
  • Hands-on experience with AI evaluation, experimentation, testing, and quality frameworks.
  • Experience using multiple LLM and agent evaluation frameworks, such as LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, or comparable platforms.
  • Ability to select, customize, and integrate evaluation frameworks for offline testing, online monitoring, regression analysis, experimentation, model and prompt comparison, and release gating.
  • Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling, and context engineering.
  • Experience evaluating non-deterministic, multi-step, or multi-agent AI systems.
  • Ability to translate human and domain-expert judgment into test cases, evaluation rubrics, scoring functions, and automated graders.
  • Experience with LLM-as-a-judge techniques, including grader design, calibration, reliability measurement, and alignment with expert human judgment.
  • Experience with regression testing, CI/CD, production monitoring, behavioral drift detection, and failure analysis.
  • Strong analytical, systems-thinking, and communication skills.

Residents of California, click Here to view our California Applicant Privacy Notice.

Anticipated Application Close Date: 11/02/26

LogicMonitor is an Equal Opportunity Employer
At LogicMonitor, we believe that innovation thrives when every voice is heard and each individual is empowered to bring their unique perspective. We’re committed to creating a workplace where diversity is celebrated, and all employees feel inspired and supported to contribute their best.

For us, equal opportunity means fostering a truly inclusive culture where everyone has the chance to grow and succeed. We don’t just open doors; we invite you to step through and be part of something bigger. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Work Authorization:
At this time, we are able to consider candidates who are authorized to work in the United States on a full-time, permanent basis without requiring new or initial employer-sponsored work authorization.
Candidates who currently hold valid U.S. work authorization that can be transferred to a new employer (such as certain H-1B statuses) may be considered on a case-by-case basis.
We are not able to provide new sponsorship for employment-based visas that require an initial petition or application by the employer.

#LI-JP1 #LI-Hybrid #BI-Hybrid

LogicMonitor is dedicated to fostering a culture of transparency and fairness, including our commitment to pay transparency. We provide the base salary ranges for all positions posted within the United States. 

Compensation packages at LogicMonitor for eligible roles include base salary, a variable plan depending on role, along with comprehensive benefits. The range displayed on each job posting reflects the minimum and maximum base salary target for new hires in the position, determined by work location and additional factors, including job-related skills, experience, interview performance, and relevant education or training. As part of our holistic compensation philosophy, your package will also include, but is not limited to: Comprehensive health, dental and vision coverage, generous parental leave policies, access to our Employee Assistance Program and various Wellness programs, a 401K with company matching, a Lifestyle Spending Account, and an unlimited vacation policy. For more information on our benefits, see our careers page.

The Base Salary range for this role is:
$158,400$217,800 USD

                                               

Our goal is to ensure an accessible and inclusive experience for every candidate.

If you need a reasonable accommodation during the application or interview process under applicable local law, please submit a request via this Accommodation Request Form.

Know your rights: workplace discrimination is illegal. Please click here to review LogicMonitor’s U.S. Pay Transparency Nondiscrimination Provision.

Skills

PythonCI/CDMachine Learning

Similar Jobs

30

Sr. Evaluation Engineer

Analogdevices · India, Bangalore, Nova

3 months ago

Senior Rust Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)

G2I · Miami +35 · Remote

4 days ago

Senior Python Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)

G2I · Miami +35 · Remote

4 days ago

Senior Go Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)

G2I · Miami +35 · Remote

4 days ago

Senior Engineer - Design & Technology Evaluation (DTE) Group

Tyndall National Institute · Cork, Cork

5 days ago

Senior Software Engineer, Agent Simulation and Evaluation

Nvidia · US, CA, Santa Clara, United States of America +1 · Remote

6 days ago

Senior Software Engineer, Agent Simulation and Evaluation

Nvidia · Santa Clara, CA,US, US +1 · Remote

6 days ago

Senior Software Engineer - Scientific Evaluation

Nvidia · US, CA, Santa Clara, United States of America

1 week ago

Senior Software Engineer - Scientific Evaluation

Nvidia · Santa Clara, CA,US, US

1 week ago

Senior Test and Evaluation Engineer

KBR Careers · USA, Colorado Springs, Schriever SFB, 210 Falcon Pkwy, Unit 156, Colorado Springs, CO, 80912, Colorado, United States of America

1 week ago

Senior Applied AI Evaluation Engineer - RAN Ticket Intelligence

Parallelwireless · Kfar Saba · Onsite

2 weeks ago

Senior RAN Systems Engineer- POC & Technical Evaluation

Parallelwireless · Kfar Saba · Onsite

2 weeks ago

Senior Reservoir Engineer - Asset Evaluation

Devonenergy · Houston, United States of America · Onsite

2 weeks ago

Senior Software Engineer, Simulation Evaluation

General Motors · GM Automation - Sunnyvale - GM Automation - Sunnyvale, United States of America · Hybrid

2 weeks ago

Senior Test & Evaluation Engineer, Connected Warfare

Andurilindustries · Aberdeen, Maryland, United States +1

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · London, England, United Kingdom · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Ottawa, ON, Canada · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Warsaw, Masovian Voivodeship, Poland · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Jakarta, Jakarta, Indonesia · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Phú Mỹ, Ho Chi Minh, Vietnam · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · New Delhi, DL, India · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Cebu City, Central Visayas, Philippines · Remote

3 weeks ago

Copy of Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · São Paulo, SP, Brazil · Remote

3 weeks ago

Senior Lua Developer (Roblox) – AI Code Evaluation

Lifted · Chicago, IL, United States · Remote

3 weeks ago

Senior Software Engineer - Model Evaluation & AI Systems

Deepgram · USA | Remote · Remote

3 weeks ago

Senior Software Systems Engineer - Simulation Evaluation & Validation

Zoox · Foster City, CA +1 · Hybrid

3 weeks ago

Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles

Nvidia · Santa Clara, CA,US, US +4 · Remote

3 weeks ago

Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles

Nvidia · US, CA, Santa Clara, United States of America +4 · Remote

3 weeks ago

Senior Test & Evaluation Engineer

Dodmg · Malibu, CA · Onsite

3 weeks ago

Senior Project Engineer, Test and Evaluation

Andurilindustries · Costa Mesa, California, United States; San Clemente, California, United States +1

3 weeks ago