- Salary
- $175k – $275k
- Workplace
- Remote, Onsite
- Type
- Full-time
- Department
- Engineering
- Experience
- 2+ years
- Source
- RecruiterFlow
Description
AI Engineer, RL & Evals
Location
San Francisco, CA
On-site — 5 days per week in-office.
Company Stage of Funding
Seed Stage / Well-Funded AI Startup
Office Type
On-site — 5 days per week in-office
Salary
$175,000 – $275,000 Base
Flexibility to go higher for exceptional candidates.
Equity
Competitive Equity
Visa
Open to Visa Transfers, including OPT and H-1B transfers. Other visa situations may be considered case-by-case.
Experience
2–7+ years of experience as an AI Engineer, Machine Learning Engineer, Reinforcement Learning Engineer, or strong backend/full-stack engineer with significant RL, evals, or ML post-training experience.
Employment Type
Full-time
Hiring Count
2 candidates
Company Description
This is a fast-growing, seed-stage AI company building the data and infrastructure layer that helps improve AI model performance in subjective and difficult-to-evaluate domains.
The company works directly with leading AI research organizations and application-layer companies on post-training, reinforcement learning environments, evaluation systems, and high-quality data infrastructure.
The engineering team operates at the intersection of backend engineering, applied machine learning, reinforcement learning, and product development. Engineers have significant ownership over the systems and environments that help evaluate and improve modern AI models.
This is not a pure research role. The ideal candidate is a product-oriented AI engineer who enjoys building production systems, shipping end-to-end software, and working hands-on with RL environments, evaluation frameworks, agent systems, and ML infrastructure.
You will work on technically challenging problems where traditional automated evaluation is difficult, including creative and subjective domains where quality cannot always be measured through simple deterministic metrics.
The ideal candidate is someone who can move comfortably between backend engineering and applied ML, take ownership of ambiguous problems, and turn research concepts into reliable production systems.
What You Will Do
1. Build & Scale RL Environments and Evaluation Systems
- Build and scale reinforcement learning environments for complex and subjective domains.
- Design evaluation frameworks that measure model performance beyond traditional benchmark metrics.
- Create unique tasks, grading systems, and evaluation methodologies for difficult-to-verify domains.
- Build systems that allow AI models and agents to be tested systematically.
- Develop infrastructure for evaluating model behavior, quality, reliability, and performance.
- Design automated evaluation workflows that reduce reliance on manual human assessment.
- Iterate on environments and evaluation systems based on model performance and research findings.
- Build reliable production infrastructure supporting RL and post-training workflows.
2. Build Backend & AI Infrastructure
- Design and build production backend systems supporting AI and ML workflows.
- Build agent harnesses, context layers, APIs, and supporting infrastructure.
- Develop data pipelines for scraping, indexing, embedding, processing, and evaluating large datasets.
- Build distributed systems capable of supporting large-scale ML and evaluation workloads.
- Develop infrastructure connecting models, environments, datasets, agents, and evaluation systems.
- Work across backend engineering and ML systems to turn ideas into production-ready products.
- Improve system scalability, reliability, performance, and maintainability.
- Own backend and infrastructure components from architecture through production.
3. Ship Product-Oriented AI Systems
- Ship production code end-to-end rather than working exclusively on research prototypes.
- Build AI-powered products and infrastructure used by internal teams and external partners.
- Develop systems that improve the quality and usefulness of AI models in real-world applications.
- Build agentic systems and evaluation infrastructure for environments where traditional automated metrics are insufficient.
- Translate ambiguous product and research requirements into practical engineering solutions.
- Work across product, engineering, and research requirements to deliver production systems.
- Rapidly prototype, validate, and productionize new ideas.
- Balance technical experimentation with reliability and production quality.
4. Collaborate With Research & Frontier AI Teams
- Collaborate with internal research teams on RL, post-training, evaluation, and model improvement.
- Work directly with leading AI labs and technical partners to develop environments and evaluation frameworks.
- Translate research concepts into production engineering systems.
- Help define tasks, environments, evaluation methodologies, and technical requirements.
- Communicate technical decisions and tradeoffs clearly across engineering and research teams.
- Contribute to technical strategy around AI evaluation and post-training infrastructure.
- Work across backend, ML, data, and product teams to solve complex AI problems.
- Take significant ownership over systems that directly influence model performance.
Ideal Candidate Background
Experience Requirements
- 2+ years of professional experience as an AI Engineer, ML Engineer, RL Engineer, or strong software engineer working on AI systems.
- Strong production engineering experience.
- Experience building reinforcement learning environments, evaluation systems, or ML post-training infrastructure.
- Experience shipping production software rather than working exclusively on research.
- Strong backend or full-stack engineering experience.
- Experience owning technical projects from design through production.
- Experience working with ambiguous technical problems.
- Experience working in fast-paced, engineering-driven environments.
- Strong ability to bridge software engineering and applied ML.
- Experience collaborating with research, product, or engineering teams.
Technical Requirements
- Strong Python experience.
- Strong PyTorch experience.
- Hands-on reinforcement learning experience.
- Experience building ML evaluation frameworks or evaluation systems.
- Experience working with LLMs.
- Strong backend engineering fundamentals.
- Experience with distributed systems.
- Experience building data pipelines.
- Experience with APIs and production services.
- Experience with data processing, indexing, embeddings, or retrieval systems.
- Strong testing and debugging practices.
- Experience building production ML or AI infrastructure.
RL, Evals & AI Systems Requirements
- Experience building reinforcement learning environments.
- Experience designing or implementing evaluation frameworks.
- Understanding of RL concepts and practical application.
- Experience evaluating LLM or agent behavior.
- Experience building systems for model post-training.
- Experience designing tasks, benchmarks, or grading systems.
- Experience working with agentic systems or agent frameworks is a strong plus.
- Experience building agent harnesses or context layers is a strong plus.
- Comfortable working on domains where objective evaluation is difficult.
- Strong understanding of the interaction between models, data, environments, and evaluation.
- Experience turning ML research concepts into production systems.
- Ability to work across model-facing and software infrastructure layers.
Soft Skills
- High ownership and accountability.
- Strong technical judgment.
- Product-oriented mindset.
- Comfortable operating in ambiguity.
- Strong problem-solving and systems-thinking skills.
- Able to move quickly from prototype to production.
- Strong communication skills.
- Comfortable collaborating with research and engineering teams.
- Able to explain technical concepts clearly.
- Strong English communication skills.
- Curious and motivated by difficult AI problems.
- Comfortable working in a fast-paced startup environment.
- Willing to work on-site 5 days per week in San Francisco.
Compensation & Benefits
- $175,000 – $275,000 base salary.
- Flexibility to go higher for exceptional candidates.
- Competitive equity.
- Full-time position.
- On-site work model — 5 days per week in San Francisco.
- Open to visa transfers, including OPT and H-1B transfers.
- Opportunity to work directly on RL, AI evaluation, post-training, and AI infrastructure.
- Significant ownership over production AI systems.
- Opportunity to work with leading AI research organizations and technical partners.
- Opportunity to build infrastructure for difficult and emerging AI evaluation problems.
Why Join
- Work at the intersection of AI engineering, reinforcement learning, evaluation, and product development.
- Build the infrastructure that helps improve the performance of modern AI models.
- Work directly with leading AI labs and advanced AI teams.
- Solve difficult evaluation problems in subjective and creative domains.
- Own RL environments, evaluation frameworks, agent systems, and backend infrastructure end-to-end.
- Work in a fast-growing seed-stage company with significant engineering ownership.
- Combine software engineering with hands-on applied ML work.
- Ship production systems rather than working exclusively on research prototypes.
- Work on technically challenging problems at the frontier of AI development.
- Help establish the technical foundation for a rapidly scaling AI company.