Hiring.Camp

ML Infrastructure Engineer

RFS Group

·

Yesterday

Salary
$200k – $300k
Workplace
Remote, Onsite
Type
Full-time
Department
Engineering
Education
Bachelor
Visa
Sponsored
Source
RecruiterFlow

Description

 
Recruiting from Scratch is a premier talent firm that focuses on placing the best product managers, software, and hardware talent at innovative companies. Our team is 100% remote and we work with teams across the United States to help them hire.

ML Infrastructure Engineer

Location - San Francisco, CA - (On-site – 5 Days Per Week)

Compensation - $200,000 – $300,000 Base + Competitive Equity

Visa - Visa Sponsorship Available (New H1B, H1B Transfers, TN, Eligible Visa Types)

Company Stage - Series B ($100M+ Raised)

Industry

Artificial Intelligence, Machine Learning Infrastructure, AI Infrastructure, Developer Tools, Enterprise AI, Document AI, Enterprise SaaS


About the Company

The company is building one of the leading AI infrastructure platforms powering document intelligence for modern AI applications.

Its platform transforms complex documents—including PDFs, spreadsheets, presentations, images, and enterprise files—into structured, production-ready data for AI systems. Processing tens of millions of pages every month, the platform serves hundreds of customers ranging from AI-native startups to Fortune 10 enterprises, enabling reliable document ingestion for large language models and AI workflows.

Backed by Andreessen Horowitz, Benchmark, and First Round Capital with over $100M in funding, the company has experienced exceptional growth, increasing revenue more than 8x year-over-year while becoming one of the fastest-growing AI infrastructure companies in the market.

As an ML Infrastructure Engineer, you'll own critical training and inference infrastructure powering production machine learning systems. You'll work closely with ML researchers and product engineers to ensure models are trained, deployed, monitored, and served efficiently while continuously improving latency, reliability, scalability, and cost across production AI workloads.

This is an exceptional opportunity to join a fast-growing AI-native company where ML Infrastructure Engineers build the foundational systems enabling next-generation AI products at massive production scale.


What You'll Do

  • Build and maintain production ML model serving infrastructure powering millions of document processing requests
  • Design and improve training infrastructure supporting models ranging from hundreds of millions to tens of billions of parameters
  • Optimize inference latency, throughput, reliability, and GPU utilization across production systems
  • Develop observability, monitoring, logging, and alerting across the ML infrastructure stack
  • Build internal tooling and data pipelines enabling ML researchers to rapidly deploy models into production
  • Architect infrastructure that intelligently routes inference workloads across multiple cloud providers
  • Optimize infrastructure for model accuracy, latency, reliability, and operational cost
  • Collaborate closely with ML researchers, software engineers, and product teams to accelerate experimentation
  • Improve Kubernetes-based infrastructure supporting large-scale ML workloads
  • Debug complex production infrastructure issues involving GPUs, distributed systems, and model serving
  • Design scalable systems supporting rapid AI model deployment and iteration
  • Champion engineering excellence across ML infrastructure, automation, and production reliability

Ideal Candidate Background

Experience Requirements

  • 3+ years of ML Infrastructure Engineering experience
  • Experience building production ML model serving and inference infrastructure
  • Experience working at AI-native startups or organizations training production ML models
  • Experience deploying and maintaining large-scale production ML systems
  • Experience supporting ML researchers with production infrastructure
  • Strong startup ownership mentality with demonstrated engineering execution
  • Experience building highly scalable infrastructure supporting AI products
  • Experience collaborating closely with ML, Product, and Infrastructure teams
  • Experience optimizing production inference systems
  • Previous experience working with GPU-intensive workloads strongly preferred

Technical Requirements

  • Strong software engineering fundamentals across distributed systems and infrastructure
  • Expert-level Python programming skills
  • Strong Kubernetes, Docker, Helm, and cloud infrastructure experience
  • Deep understanding of GPU optimization, model serving, and inference infrastructure
  • Experience with PyTorch and modern ML deployment workflows
  • Experience building observability, monitoring, and logging systems for ML infrastructure
  • Strong understanding of model deployment, training pipelines, and production inference
  • Experience working with 1–3 node model training and single/double-node serving
  • Familiarity with cloud platforms including AWS, Azure, or GCP
  • Ability to rapidly debug production infrastructure and performance bottlenecks

Education

  • Bachelor's degree in Computer Science, Machine Learning, Engineering, Mathematics, or another technical discipline preferred
  • Strong engineering background with demonstrated infrastructure expertise
  • Top technical universities, competitive programming backgrounds, or exceptional engineering accomplishments strongly preferred
  • Publications at top ML conferences (NeurIPS, ICML, ICLR, CVPR) are a strong plus

Soft Skills

  • Outstanding technical problem-solving
  • Strong engineering ownership
  • Excellent collaboration skills
  • High execution mentality
  • Comfortable operating with ambiguity
  • Strong cross-functional communication
  • Startup mindset with high execution velocity
  • Structured systems thinker
  • Passion for AI infrastructure and production ML systems
  • Continuous learner with deep technical curiosity

Compensation & Benefits

  • Base Salary: $200,000 – $300,000
  • Competitive Equity Package
  • Visa Sponsorship Available
  • On-site work model (5 days/week in San Francisco)
  • Opportunity to build core infrastructure powering one of the fastest-growing AI platforms
  • Direct collaboration with ML researchers and engineering leadership
  • Significant ownership across production AI infrastructure
  • High-impact engineering culture with exceptional career growth
  • Well-funded Series B company backed by leading venture firms
  • Opportunity to build foundational infrastructure enabling next-generation AI systems

Why Join

This is an opportunity to build the infrastructure powering one of the fastest-growing AI platforms transforming how enterprises process information.

You'll work at the intersection of machine learning infrastructure, distributed systems, GPU optimization, model serving, production engineering, and AI infrastructure while solving some of the hardest scalability challenges in modern AI.

As an ML Infrastructure Engineer, you'll own foundational systems that enable production AI models to operate reliably, efficiently, and at massive scale while helping define the future of enterprise AI infrastructure.

 

Skills

PythonAWSAzureGCPDockerKubernetesMachine LearningPyTorch

Similar Jobs

30

ML Infrastructure Engineer

xAI · Palo Alto, California, United States

1 week ago

ML Infrastructure Engineer

White Circle · Paris or London · Hybrid

4 weeks ago

ML Infrastructure Engineer

Mach9 · San Francisco, US

1 month ago

ML Infrastructure Engineer

techire ai · San Francisco, CA · Hybrid, Onsite

3 months ago

ML Infrastructure Engineer

Mach9 · San Francisco · Onsite

3 months ago

ML Infrastructure Engineer

Applovin · Palo Alto, CA

5 months ago

ML Infrastructure Engineer

Sunday · Mountain View · Remote

5 months ago

ML Infrastructure Engineer

Echo Neurotechnologies · San Francisco · Hybrid

6 months ago

Software Engineer - ML Infrastructure

Andurilindustries · Costa Mesa, California, United States

3 days ago

ML Infrastructure Engineer, Fauna

Amazon · Remote

1 week ago

Software Engineer, ML Infrastructure

Ideogram · Toronto · Onsite

3 weeks ago

Staff ML Infrastructure Engineer

General Motors · GM Automation - Sunnyvale - GM Automation - Sunnyvale, United States of America · Remote

1 month ago

Software Engineer, ML Infrastructure, Level 4

Snapchat · Santa Monica - 3100 Ocean Park Blvd, United States of America +2

1 month ago

Staff Software Engineer, ML Infrastructure, Level 6

Snapchat · Bellevue - 110 110th Ave NE, United States of America +2

1 month ago

HPC/ML Infrastructure Engineer

Spellbrush · San Francisco or Tokyo · Onsite

1 month ago

Data Infrastructure & ML Engineer (Hybrid Role)

Axcelis Technologies has been at · Beverly MA, United States of America · Remote, Hybrid

1 month ago

Senior ML Infrastructure Engineer

Arlo · New York City · Onsite

2 months ago

Staff Software Engineer, ML Infrastructure

Simplisafe · Boston, MA +1 · Remote, Hybrid

2 months ago

Staff Machine Learning Engineer, ML Infrastructure

Simplisafe · Boston, MA +1 · Hybrid

2 months ago

Staff Software Engineer, ML Infrastructure

Voxel · San Francisco, CA · Hybrid

2 months ago

Senior/Staff Software Engineer - ML Infrastructure

Voxel · San Francisco, CA · Hybrid

3 months ago

Staff ML Infrastructure Engineer (GPU & Distributed Systems)

techire ai · Remote, Hybrid

3 months ago

Staff AI/ML Infrastructure Engineer

Vultr · Remote - United States · Remote

3 months ago

Data/ML Infrastructure Engineer

Matter Intelligence · San Francisco · Onsite

4 months ago

Staff ML Infrastructure Engineer - Embodied AI

General Motors · Sunnyvale Technical Center - Sunnyvale Technical Center (CL), United States of America +1 · Remote

4 months ago

Senior ML Infrastructure Engineer - Embodied AI

General Motors · Sunnyvale Technical Center - Sunnyvale Technical Center (CL), United States of America +1 · Remote

4 months ago

Senior ML Infrastructure Engineer

Prior Labs · Freiburg or Berlin

4 months ago

Staff+ Data Engineer (ML Infrastructure)

Sanas · Palo Alto, CA

4 months ago

Cloud and ML Infrastructure Engineer

Glimpse · Somerville · Hybrid

4 months ago

Senior ML Infrastructure Engineer - Embodied AI Scaling Foundations

General Motors · GM Automation - Sunnyvale - GM Automation - Sunnyvale, United States of America +1 · Hybrid

4 months ago