Machine Learning Engineer II, GenAI Evaluation Framework
Cox brings together the world
·Today
- Salary
- $92k – $154k
- Location
- Atlanta, GA - 6305 Peachtree Dunwoody Rd Bldg B, United States of America
- Workplace
- Remote, Hybrid
- Type
- Full-time
- Department
- Engineering
- Visa
- Not sponsored
- Closing date
- Today
- Source
- Workday
Description
Company
Cox Automotive - USAJob Family Group
Job Profile
Management Level
Flexible Work Option
Travel %
Work Shift
Compensation
Compensation includes a base salary in the range of $92,300.00 - $153,900.00. The base salary may vary within the anticipated base pay range based on factors such as the ultimate location of the position and the selected candidate’s knowledge, skills, and abilities. Position may be eligible for additional compensation that may include an incentive program.Job Description
About the AI Accelerator
The AI Accelerator drives AI innovation across Cox Automotive. The team builds proof-of-concept agents, ships reusable AI components, sets evaluation and governance standards, and partners with product teams to bring AI into production.
About the Role
Cox Automotive is hiring a Machine Learning Engineer II to own and build the GenAI Evaluation Framework (GEF), the shared capability that measures whether AI agents running in production meet quality standards. This is a builder-owner role. You own GEF as an internal product, which means you set the roadmap, decide what ships and what gets retired, and sequence which teams onboard. You also write the evaluators, pipelines, and reference implementations yourself.
The framework already exists and has been proved in one production implementation with the Virtual Contact Assistant. The current work is turning that implementation into a product other teams adopt on its merits, since adoption across business units is voluntary. You will have the direct support of the Director of Machine Learning and AI on prioritization and executive alignment, and full latitude on the technical and product decisions inside GEF.
What You Will Own
- GEF product roadmap. Shared accountability with AI leaders of what ships next, what to generalize from the first production implementation, and what to deprecate. You help to maintain the roadmap and the reasoning behind it.
- Enablement queue. You sequence which product teams onboard, when, and against what success criteria, and you negotiate that sequence with the teams involved.
- Framework build. You implement and maintain LLM-as-judge, rule-based, and statistical evaluators, and calibrate them against human-labeled baselines.
- Evaluation pipelines. You build and operate the pipelines that score agent traffic on a schedule and publish results to dashboards and the Cox Agent Registry.
- Golden datasets. You work with product teams and subject matter experts to assemble, version, and maintain the datasets evaluation depends on.
- Onboarding and adoption. You integrate GEF into adopting teams' agents, debug failures, and write the documentation and reference implementations that let the next team self-serve.
- Quality reporting. You publish framework adoption, agent quality health, and cost-per-quality metrics each month, and supply the audit trails and evaluation evidence governance reviews require.
- Agent instrumentation. You help add tracing and logging to agent workflows and integrate evaluation checks into CI and deployment pipelines.
What Success Looks Like in Year One
- Three to four product teams onboarded to GEF, with golden datasets curated and evaluators calibrated against human baselines.
- A published GEF roadmap that adopting teams and leadership both plan against.
- Framework capability a new team can adopt from documentation and reference implementations without direct support.
- Quality health published per registered agent in the Cox Agent Registry, on a monthly cadence, without manual assembly.
- Pipelines that run reliably, with monitoring and runbooks that let someone else operate them.
Who You Will Work With
Inside the AI Accelerator: the engineers building GEF, the team embedded with the Virtual Contact Assistant pilot, and the broader CADS data science group.
Across Cox Automotive: engineering and product teams at VinSolutions, Autotrader, Manheim, Dealertrack, and other business units as they onboard, plus Platform Engineering, Architecture, Enterprise AI, Security, and FinOps.
Required Qualifications
- Two or more years of professional machine learning engineering experience with a Bachelor’s Degree, OR up to two years combined with a Master’s Degree in a quantitative field.
- Applicants must currently be authorized to work in the United States for any employer without current or future sponsorship. No OPT, CPT, STEM/OPT or visa sponsorship now or in future.
- Must currently live within a commutable distance to Atlanta to be considered.
- Experience owning the direction of a technical product, internal tool, library, or platform used by other teams. Formal product ownership is welcome, and so is having been the de facto owner of something other engineers depended on.
- Ability to maintain a roadmap, set priorities among competing requests, and explain the tradeoffs to both engineers and product stakeholders.
- Strong Python, with working habits around testing, version control, code review, and continuous integration.
- Hands-on experience building with large language models, including prompting, API integration, retrieval, or agent frameworks.
- Working knowledge of measurement concepts such as sampling, inter-rater agreement, precision and recall, or comparative testing.
- Experience with cloud services, ideally AWS, and with building data pipelines.
- Strong written communication, including documentation, roadmap updates, and reporting to senior stakeholders.
Preferred Qualifications
- Formal product ownership or technical product management experience for developer tools, infrastructure, or ML platforms.
- Experience with evaluation tooling such as DeepEval, AWS Bedrock AgentCore Evaluations, Langfuse, or RAGAS.
- Experience in a federated organization where adoption is voluntary and the product must win on merit.
- Experience with observability and tracing tools such as OpenTelemetry or Datadog.
- Exposure to model risk, AI governance, or work in a regulated environment.
- Prior work at the intersection of data science teams and product or platform engineering.
Drug Testing
Benefits
About Us