Hiring.Camp

Research Engineer

RFS Group

·

Today

Salary
$250k – $290k
Location
San Francisco, California
Workplace
Remote, Hybrid
Type
Full-time
Department
Engineering
Experience
3+ years
Education
Bachelor
Source
RecruiterFlow

Description

 
Recruiting from Scratch is a premier talent firm that focuses on placing the best product managers, software, and hardware talent at innovative companies. Our team is 100% remote and we work with teams across the United States to help them hire.

Research Engineer

Location

San Francisco, CA / Bay Area, CA

Hybrid role. Remote-friendly culture with an SF hub; candidates in the Bay Area are preferred for hybrid collaboration.

Compensation

$250,000 – $290,000 Base + Competitive Equity

Visa

Open to visa transfers, including H-1B and OPT transfers.

Exceptional candidates may be considered for additional sponsorship on a case-by-case basis.

Company Stage

High-Growth / Venture-Backed Technology Company

Industry

Artificial Intelligence, AI Infrastructure, Developer Tools, Machine Learning, LLMs, Web Data, Data Infrastructure, Web Scraping, Developer Platforms


About the Company

Our client is building AI infrastructure that enables developers and AI-native companies to reliably turn the web into clean, structured, LLM-ready data.

The platform processes web content at massive scale, converting URLs into structured data and markdown that can be consumed by LLM applications, AI agents, search systems, and modern AI workflows.

The company has experienced exceptional growth, reaching eight figures in ARR during its first year and continuing to scale rapidly. Its developer-focused product has become widely adopted across the AI ecosystem, with a large open-source community and thousands of developers and AI companies building on the platform.

This is a small, highly technical, high-ownership engineering organization where engineers own meaningful pieces of the product end-to-end.

The Research Engineer role sits at the intersection of machine learning, research methodology, backend engineering, data engineering, and evaluation.

This is not a traditional research role where the primary responsibility is publishing papers or experimenting with models in isolation. It is also not a benchmark-operator role where someone simply runs pre-existing evaluation frameworks.

Instead, the engineer will define what "good" means, design the metrics used to measure it, build the evaluation infrastructure, generate and curate datasets, and create feedback loops that directly influence model and product decisions.

The ideal candidate combines strong engineering ability with rigorous analytical thinking and is comfortable working in ambiguous environments where the evaluation methodology itself needs to be invented.


What You'll Do

  • Design evaluation systems that rigorously measure the quality of AI and web-data outputs
  • Define metrics that determine what "good output" means across millions of websites and edge cases
  • Build scalable evaluation pipelines and testing infrastructure
  • Build evaluation harnesses capable of measuring model and product performance at scale
  • Generate and curate high-quality datasets for evaluation
  • Design datasets that accurately represent real-world web diversity and edge cases
  • Develop methodologies for evaluating LLM-powered systems
  • Measure the accuracy, reliability, completeness, and usefulness of generated outputs
  • Build automated evaluation systems for web extraction and structured data generation
  • Create statistical and quantitative methods for comparing system performance
  • Investigate failure cases and identify patterns in model or product behavior
  • Develop feedback loops connecting evaluation results to model improvements
  • Translate evaluation findings into actionable product and engineering decisions
  • Work closely with ML engineers and researchers to improve model performance
  • Partner with product and engineering teams to determine evaluation priorities
  • Build tooling that makes evaluation results accessible and actionable across the company
  • Develop systems for continuously monitoring output quality
  • Create regression testing frameworks for AI-powered systems
  • Establish benchmarks and quality standards for production AI systems
  • Build data pipelines supporting evaluation workflows
  • Automate dataset generation, labeling, validation, and analysis
  • Work with large-scale web and text datasets
  • Investigate edge cases across different websites, formats, languages, and content types
  • Analyze NLP and LLM outputs for quality and consistency
  • Develop methods for evaluating structured and unstructured outputs
  • Design experiments to determine whether product and model changes actually improve performance
  • Analyze statistical significance and reliability of evaluation results
  • Identify limitations and biases within evaluation datasets
  • Improve evaluation methodology as models and product capabilities evolve
  • Build internal tools for researchers and engineers to investigate system behavior
  • Own evaluation infrastructure from initial concept through production deployment
  • Move quickly from hypothesis to implementation, measurement, and iteration
  • Close the loop between research, engineering, evaluation, and product decisions
  • Work on ambiguous problems where there may not be an established methodology
  • Contribute to technical direction around AI evaluation and measurement
  • Operate with high autonomy in a fast-moving AI infrastructure environment
  • Work directly with founders and senior technical leadership on high-impact problems
  • Build systems that help the company understand whether its core product actually works

Ideal Candidate Background

Experience Requirements

  • 3+ years of experience in Machine Learning Engineering, Research Engineering, Software Engineering, Data Engineering, Applied Science, or related technical roles
  • Strong engineering foundation with experience building production systems
  • Experience building evaluation systems, testing infrastructure, experimentation frameworks, or data-heavy systems
  • Experience working with machine learning or LLM-powered systems
  • Experience designing metrics or methodologies for measuring system quality
  • Experience building data pipelines or infrastructure supporting evaluation and experimentation
  • Experience working with large datasets
  • Experience analyzing complex or noisy data
  • Experience solving ambiguous technical problems independently
  • Experience taking research or analytical concepts and turning them into production systems
  • Experience designing experiments and interpreting results
  • Strong understanding of how to measure model or system performance
  • Experience working with NLP, LLMs, web data, search, recommendation systems, or similar domains is highly relevant
  • Experience working across engineering and research-oriented teams
  • Strong ability to balance methodological rigor with practical engineering constraints
  • Comfortable owning projects end-to-end
  • Comfortable operating without fully specified requirements
  • Strong bias toward shipping, measuring, and iterating
  • Experience working in startups or high-growth technology environments strongly preferred

Technical Requirements

  • Strong Python engineering skills
  • Strong software engineering fundamentals
  • Experience building production-grade data pipelines
  • Experience working with machine learning systems
  • Experience working with LLMs or NLP systems
  • Experience designing evaluation frameworks
  • Experience creating quantitative metrics
  • Strong statistical analysis fundamentals
  • Strong experimental design skills
  • Experience analyzing model outputs
  • Experience evaluating AI-generated content
  • Experience working with large-scale datasets
  • Experience with data generation and dataset curation
  • Experience with data cleaning and preprocessing
  • Experience building automated testing or evaluation systems
  • Experience designing regression testing frameworks
  • Experience building scalable evaluation infrastructure
  • Experience with machine learning experimentation
  • Experience with statistical comparison of model or system performance
  • Ability to identify meaningful evaluation signals from noisy data
  • Ability to distinguish useful metrics from misleading metrics
  • Ability to reason about evaluation methodology and measurement bias
  • Strong understanding of data quality
  • Experience with NLP datasets is a plus
  • Experience with web data or web scraping is a plus
  • Experience with information retrieval is a plus
  • Experience with search evaluation is a plus
  • Experience with structured data extraction is a plus
  • Experience with LLM evaluation is a strong plus
  • Experience with RAG evaluation is a plus
  • Experience with agent evaluation is a plus
  • Experience with benchmark design is a plus
  • Experience with experiment tracking and analysis is a plus
  • Experience with distributed systems is a plus
  • Experience with statistical or scientific computing libraries is a plus
  • Experience with PyTorch, TensorFlow, or similar ML frameworks is a plus
  • Experience with modern LLM evaluation frameworks is a plus
  • Ability to write maintainable, production-quality code
  • Ability to debug complex data and evaluation pipelines
  • Ability to reason about system performance at scale
  • Ability to communicate technical and methodological tradeoffs clearly

Education

  • Bachelor's degree in Computer Science, Mathematics, Statistics, Engineering, Machine Learning, Data Science, or a related technical discipline preferred
  • Advanced degree in Computer Science, Machine Learning, Statistics, NLP, or a related field is a plus
  • Strong quantitative or technical academic background preferred
  • Research experience can be valuable, provided the candidate also has strong engineering execution
  • Strong professional engineering experience can compensate for academic pedigree

Soft Skills

  • Strong first-principles problem solver
  • Extremely analytical
  • Methodologically rigorous
  • Strong intellectual curiosity
  • Comfortable defining problems from scratch
  • Comfortable inventing new metrics and evaluation methodologies
  • Strong engineering judgment
  • Strong research instincts
  • Strong product thinking
  • Comfortable working with ambiguity
  • High agency
  • Strong ownership mentality
  • Bias toward action
  • Fast learner
  • Comfortable experimenting and iterating
  • Willing to challenge assumptions
  • Strong attention to detail
  • Comfortable investigating unexpected system behavior
  • Able to distinguish signal from noise
  • Comfortable working with imperfect datasets
  • Strong quantitative reasoning
  • Able to communicate complex findings clearly
  • Comfortable collaborating with researchers, ML engineers, and product engineers
  • Able to translate analytical findings into engineering decisions
  • Comfortable operating in high-velocity startup environments
  • Motivated by measurable real-world impact
  • Comfortable shipping systems rather than staying in theoretical research
  • Comfortable working without a fully defined playbook
  • Strong ability to close feedback loops
  • Motivated by understanding whether systems actually work

Compensation & Benefits

  • Base Salary: $250,000 – $290,000
  • Competitive equity package
  • Opportunity to work on cutting-edge AI infrastructure
  • Opportunity to build evaluation systems used across large-scale AI workloads
  • Direct impact on AI product quality and model decisions
  • High technical ownership and autonomy
  • Small, highly technical engineering team
  • Direct exposure to founders and senior technical leadership
  • Fast-paced AI infrastructure environment
  • Opportunity to work with LLMs, NLP, web data, and large-scale evaluation systems
  • Remote-friendly culture with an SF/Bay Area hub
  • Opportunity to work on infrastructure used by leading AI companies and developers
  • Significant opportunity for technical growth and scope

Why Join

This is an opportunity to work on one of the hardest problems in modern AI infrastructure: how do you rigorously determine whether an AI-powered system actually works?

You'll build the evaluation systems, datasets, metrics, and feedback loops that determine the quality of AI-generated web data at massive scale.

Unlike traditional research roles, you'll not only investigate methodologies — you'll build the systems that operationalize them.

Unlike traditional software engineering roles, you'll spend significant time defining metrics, designing experiments, analyzing model behavior, and determining what quality actually means.

Your work will directly influence model development, product decisions, and the reliability of infrastructure used by AI-native companies.

The role is ideal for someone who enjoys operating at the intersection of software engineering, machine learning, research methodology, data, and product.

If you enjoy ambiguous problems, building from first principles, designing rigorous measurement systems, and moving quickly from hypothesis to implementation and iteration, this role offers exceptional technical scope and impact.

Skills

PythonMachine LearningNLPTensorFlowPyTorchData ScienceData Engineering

Similar Jobs

30

Research Engineer

Quantiphi · IN KA Bengaluru, India +1 · Hybrid

4 days ago

Research Engineer

Psu · Penn State University Park, United States of America · Remote, Hybrid

1 week ago

Research Engineer

Neo4J · London

1 week ago

Research Engineer

Greptile · San Francisco · Onsite

1 week ago

Research Engineer

CNI Careers · AL Fort Novosel G Farrel Rd 6901, United States of America · Onsite

2 weeks ago

Research Engineer

Tessera Labs · San Jose Office (HQ) +1

2 weeks ago

Research Engineer

E Ink · Billerica, MA

2 weeks ago

Research Engineer

USC is · Los Angeles, CA - Health Sciences Campus, United States of America

4 weeks ago

Research Engineer

Firecrawl · San Francisco, California, US · Hybrid, Onsite

1 month ago

Research Engineer

Verisign · Reston,Virginia,United States

1 month ago

Research Engineer

The LYCRA Company · WAYNESBORO, VA

1 month ago

Research Engineer

Antimetal · HQ - NYC · Onsite

1 month ago

Research Engineer

Imc · Sydney, Australia

1 month ago

Research Engineer

Imc · Hong Kong, Hong Kong

1 month ago

Research Engineer

Decagon · San Francisco · Onsite

1 month ago

Research Engineer

Superannotate · San Francisco · Hybrid

1 month ago

Research Engineer

Awetomaton · Beavercreek, OH +1

1 month ago

Research Engineer

Console · San Francisco (On-site) · Onsite

1 month ago

Research Engineer

Helm.ai · Remote - US · Remote

2 months ago

Research Engineer

Helm.ai · Remote - Canada · Remote

2 months ago

Research Engineer

Kog AI · Paris, France · Hybrid

2 months ago

Research Engineer

Datologyai · Redwood City · Remote, Onsite

3 months ago

Research Engineer

NowSecure · Remote

3 months ago

Research Engineer

AutoDesk · AMER - United States - Massachusetts - Boston - Drydock, United States of America · Hybrid

3 months ago

Research Engineer

SID · San Francisco, California, US

3 months ago

Research Programmer

Upenn · Smilow Center for Translational, United States of America

3 months ago

Research Programmer

Upenn · Smilow Center for Translational, United States of America

3 months ago

Research Engineer

Thomson Reuters · United States of America, Eagan, Minnesota +6 · Hybrid

4 months ago

Research Engineer

"Paradromics, Inc." · Oakland, CA

4 months ago

Research Engineer

Hedra · San Francisco

4 months ago