- Salary
- $150k – $200k
- Workplace
- Remote, Onsite
- Type
- Full-time
- Department
- IT
- Seniority
- Senior
- Source
- RecruiterFlow
Description
Member of Technical Staff - Research Scientist
Location
San Francisco, CA
Fully onsite, 5 days per week.
Company Stage of Funding
Early-Stage / High-Growth AI Company
Office Type
On-site — 5 days per week
Salary
$150,000 – $200,000 Base
Equity
0.15% – 0.5% with flexibility upwards
Visa
Open to visa transfers, including OPT and H1B transfers.
Experience
0–3 years of experience in applied AI/ML, with strong emphasis on benchmarking, evaluation methodologies, language models, or related AI research and engineering.
Employment Type
Full-time
Hiring Count
Growth Hiring — 5 hires
Company Description
This is an early-stage AI company building technology that helps enterprises evaluate, validate, and deploy increasingly capable AI systems with confidence.
The company is focused on solving difficult problems around AI evaluation, benchmarking, and the reliability of language models. Its work helps determine how advanced AI systems perform across real-world enterprise use cases and provides the methodologies needed to measure and improve their quality.
The engineering and research teams work at the intersection of applied AI research and production software engineering. You will evaluate new models as they are released, build benchmarks from scratch, develop evaluation methodologies, and work closely with engineering teams to turn research ideas into scalable systems.
As a Member of Technical Staff — Research Scientist, you will have significant ownership over how AI systems are evaluated. You will work with large datasets, language models, evaluation pipelines, and automated assessment methods while collaborating with AI labs and enterprise customers.
This is a highly hands-on applied research role for someone who enjoys building practical evaluation systems, experimenting with modern AI models, and turning research ideas into production-ready methodologies. The role prioritizes applied impact over purely academic research or publishing papers for its own sake.
What You Will Do
1. Build & Develop AI Evaluation Methodologies
- Evaluate new AI and language models as they are released.
- Design and build new benchmarks from scratch.
- Develop evaluation methodologies for measuring model quality, reliability, and performance.
- Improve automated evaluation methods for generated text.
- Design experiments to understand model strengths, weaknesses, and failure modes.
- Explore new approaches to benchmarking increasingly capable AI systems.
- Translate research questions into practical evaluation frameworks.
2. Build Datasets, Benchmarks & Evaluation Infrastructure
- Construct datasets for evaluating language models and AI systems.
- Work with labelers to create high-quality evaluation datasets.
- Develop benchmark specifications, annotation processes, and quality standards.
- Write technical documentation and white papers describing methodologies and results.
- Build tooling that enables repeatable and scalable model evaluation.
- Analyze evaluation results and identify opportunities to improve benchmarks.
- Work with engineering to implement and scale evaluation systems.
3. Work Across AI Research & Engineering
- Collaborate closely with software engineers to productionize evaluation methodologies.
- Build clean, maintainable Python code for research and evaluation workflows.
- Work with deep learning frameworks such as PyTorch and TensorFlow.
- Explore language modeling, diffusion models, and other modern AI/ML techniques.
- Work with Python-based web infrastructure such as Django or Flask where relevant.
- Leverage AWS infrastructure to run and scale evaluation workloads.
- Translate experimental findings into robust technical systems.
4. Collaborate With AI Labs & Enterprise Customers
- Work with leading AI labs and enterprise customers to understand evaluation needs.
- Translate customer and research requirements into benchmark and evaluation strategies.
- Communicate technical findings clearly to both technical and non-technical stakeholders.
- Incorporate feedback from engineers, researchers, customers, and other team members.
- Participate in collaborative research and development efforts.
- Help define how advanced AI systems should be tested before deployment in enterprise environments.
Ideal Candidate Background
Experience Requirements
- 0–3 years of professional experience in applied AI/ML, research engineering, or a closely related field.
- Hands-on experience working with AI/ML systems, language models, benchmarking, or evaluation.
- Experience building practical AI/ML systems rather than exclusively conducting theoretical research.
- Experience working in collaborative engineering or research environments.
- Strong interest in evaluating and understanding the behavior of modern AI systems.
- Experience designing experiments, methodologies, or measurement frameworks is highly valuable.
- Startup, AI lab, research lab, or founding/early-stage company experience is strongly preferred.
- Candidates with strong academic backgrounds and limited industry experience can also be considered.
Technical Requirements
- Strong Python programming skills.
- Ability to write clean, maintainable, production-quality code.
- Experience with AI/ML frameworks such as PyTorch or TensorFlow.
- Understanding of machine learning fundamentals.
- Familiarity with language modeling and modern generative AI techniques.
- Experience working with datasets and data-processing pipelines.
- Strong analytical and experimental skills.
- Familiarity with benchmarking or evaluation methodologies.
- Experience with cloud infrastructure, ideally AWS.
- Familiarity with Django, Flask, or other Python-based HTTP frameworks is a plus.
- Understanding of software engineering best practices, Git, pull requests, code review, and collaborative development.
AI, LLM & Evaluation Requirements
- Hands-on experience with large language models or other modern AI systems.
- Strong interest in LLM evaluation and benchmarking.
- Experience evaluating model outputs or AI-generated content.
- Understanding of automated evaluation methodologies.
- Ability to design experiments for comparing AI models or systems.
- Familiarity with benchmark construction and dataset design.
- Understanding of model failure modes and evaluation challenges.
- Experience with language modeling, NLP, or related AI research is strongly preferred.
- Ability to move from research hypothesis to measurable evaluation methodology.
- Interest in emerging AI models and rapidly changing AI capabilities.
Research & Applied AI Requirements
- Experience conducting applied AI/ML research.
- Ability to formulate research questions and design experiments.
- Strong quantitative and analytical reasoning.
- Experience analyzing experimental results and drawing actionable conclusions.
- Ability to develop new methodologies rather than only applying existing benchmarks.
- Strong preference for applied research with practical product or engineering impact.
- Academic publications in NLP, benchmarking, or related areas are a plus but are not required.
- Strong candidates do not need to be focused on publishing papers as their primary career objective.
Soft Skills
- Strong written and verbal communication.
- Comfortable giving and receiving constructive feedback.
- Strong analytical and problem-solving mindset.
- Highly curious about emerging AI technologies.
- Comfortable working in ambiguous research environments.
- Able to communicate complex technical concepts clearly.
- Collaborative and team-oriented.
- Comfortable working closely with engineers and researchers.
- Able to balance research quality with practical execution.
- Highly self-directed and intellectually curious.
- Motivated by real-world impact rather than research publication alone.
Compensation & Benefits
- $150,000 – $200,000 base salary.
- 0.15% – 0.5% equity, with flexibility upwards.
- Meaningful early-stage equity opportunity.
- Relocation support for candidates moving to San Francisco.
- Transportation support as needed.
- Opportunity to work on cutting-edge AI evaluation problems.
- Small, highly collaborative team.
- Direct exposure to advanced AI models and systems.
- Fully onsite in San Francisco.
Why Join
- Work at the forefront of AI evaluation and benchmarking.
- Help define how advanced AI systems are measured before enterprise deployment.
- Evaluate some of the newest and most capable AI models as they are released.
- Build new benchmarks and evaluation methodologies from scratch.
- Work across applied AI research and production engineering.
- Collaborate closely with engineers, AI labs, and enterprise customers.
- Gain hands-on experience with LLMs, benchmarking, datasets, and automated evaluation.
- Work on technically challenging problems where there is no established playbook.
- Have meaningful ownership over research direction and evaluation methodology.
- Join a small team where your work can directly influence the company's technical direction.