Hiring.Camp

Machine Learning Systems Engineer

RFS Group

·

Yesterday

Salary
$200k – $300k
Workplace
Remote, Onsite
Type
Full-time
Department
Engineering
Experience
2+ years
Education
Bachelor
Source
RecruiterFlow

Description

 
Recruiting from Scratch is a premier talent firm that focuses on placing the best product managers, software, and hardware talent at innovative companies. Our team is 100% remote and we work with teams across the United States to help them hire.

Machine Learning Systems Engineer

Location - Palo Alto, CA (On-site) - Five days per week in-office in the Bay Area.

Compensation - $200,000 – $300,000 Base + Competitive Equity

Visa - Open to Visa Transfers (OPT, H1B Transfers)

Company Stage - Growth Stage – $56M Funding

Industry - Artificial Intelligence, Machine Learning, Generative AI, AI Infrastructure, Developer Infrastructure


About the Company

Our client is building a new generation of highly efficient AI models designed to dramatically improve the speed and economics of large language model inference.

The company has pioneered diffusion-based language models that generate responses in parallel rather than relying exclusively on traditional sequential token generation. This approach enables significantly faster and more efficient AI inference while maintaining competitive model quality.

The company launched one of the first commercially available diffusion-based language models in early 2025 and is now deploying large-scale AI models with Fortune 500 organizations.

The team is small, highly technical, and research-driven, with engineers working directly alongside world-class researchers and founders. The organization places a strong emphasis on technical depth, experimentation, performance optimization, and production-scale AI infrastructure.

As a Machine Learning Systems Engineer, you'll work on the infrastructure that enables large-scale model training and inference, contributing directly to systems that make advanced AI models faster, more efficient, and more reliable.

This is an opportunity to join an elite AI team where you can work at the intersection of machine learning, distributed systems, GPU infrastructure, and high-performance model serving.


What You'll Do

  • Design, build, and operate infrastructure supporting large-scale ML training and inference systems
  • Develop high-performance systems for serving and deploying large language models
  • Optimize model inference for latency, throughput, memory utilization, and cost efficiency
  • Build and maintain production ML infrastructure across GPU and cloud environments
  • Work with inference engines such as vLLM, TensorRT, ONNX Runtime, and SGLang
  • Develop and optimize GPU-accelerated ML workloads using CUDA
  • Build scalable training and inference pipelines using PyTorch and/or TensorFlow
  • Deploy and manage ML workloads across Kubernetes and containerized environments
  • Design distributed systems capable of supporting high-volume model inference
  • Improve model serving performance across different hardware and infrastructure configurations
  • Build reliable systems for model deployment, monitoring, evaluation, and production operations
  • Work closely with research teams to translate new model architectures into production systems
  • Optimize infrastructure for emerging diffusion-based language models and other generative AI architectures
  • Develop tooling and automation for ML experimentation and deployment
  • Build and maintain cloud infrastructure across AWS, Azure, or comparable environments
  • Work with Kubeflow and other ML orchestration platforms
  • Investigate performance bottlenecks across compute, networking, memory, and model-serving layers
  • Develop systems that make model training and inference faster, more efficient, and more reliable
  • Contribute to system architecture and technical strategy across the ML infrastructure stack
  • Operate with high ownership in a fast-moving, deeply technical AI environment
  • Work closely with founders and researchers on highly technical infrastructure challenges

Ideal Candidate Background

Experience Requirements

  • 2–5 years of professional experience in ML Systems Engineering, ML Infrastructure, AI Infrastructure, or related engineering roles
  • Experience building and operating production ML systems
  • Experience working on infrastructure for model training and/or inference
  • Experience deploying machine learning models into production environments
  • Experience working with GPU-based computing infrastructure
  • Experience building scalable ML or distributed systems
  • Experience working with modern deep learning frameworks
  • Experience operating in technically demanding engineering environments
  • Experience collaborating closely with research and engineering teams
  • Strong ownership mentality with demonstrated execution ability
  • Comfortable working on complex technical problems with limited precedent
  • Strong interest in machine learning systems and AI infrastructure
  • Ability to operate effectively in a fast-moving, research-driven environment

Technical Requirements

  • Strong Python engineering experience
  • Strong experience with PyTorch, TensorFlow, or comparable ML frameworks
  • Experience with GPU computing and CUDA
  • Experience with ML inference systems such as vLLM, TensorRT, ONNX Runtime, or SGLang
  • Strong understanding of model serving and inference optimization
  • Experience with Docker and containerized ML workloads
  • Experience with Kubernetes
  • Experience working with AWS, Azure, or other major cloud platforms
  • Experience building distributed systems or scalable infrastructure
  • Experience with ML training and inference pipelines
  • Experience with Kubeflow or comparable ML orchestration platforms preferred
  • Strong understanding of performance optimization and system reliability
  • Experience debugging production ML infrastructure
  • Strong understanding of computer science and software engineering fundamentals
  • Ability to reason about GPU utilization, memory constraints, latency, and throughput
  • Strong debugging and performance analysis capabilities

Education

  • Bachelor's degree or higher in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, or related technical field preferred
  • Advanced degree in Machine Learning, Computer Science, or related field is a plus
  • Strong computer science, systems, and machine learning fundamentals
  • Equivalent practical engineering experience accepted

Soft Skills

  • Exceptional technical ownership
  • Strong analytical and problem-solving ability
  • Deep technical curiosity
  • Comfortable working on difficult and ambiguous infrastructure problems
  • Strong communication skills
  • Comfortable collaborating with researchers and highly technical engineers
  • High execution velocity
  • Strong attention to system performance and engineering quality
  • Bias toward experimentation and continuous improvement
  • Low-ego collaborative mentality
  • Comfortable receiving and incorporating technical feedback
  • Strong ability to reason from first principles
  • Comfortable working in a small, high-performing team
  • Strong interest in cutting-edge AI systems
  • Willingness to work five days per week in Palo Alto

Compensation & Benefits

  • Base Salary: $200,000 – $300,000
  • Competitive Equity Package
  • Opportunity to work on cutting-edge diffusion-based language models
  • Direct collaboration with world-class AI researchers and founders
  • Opportunity to work on large-scale ML training and inference infrastructure
  • Exposure to advanced GPU optimization and AI systems engineering
  • Significant technical ownership in a small, elite engineering organization
  • Opportunity to influence foundational ML infrastructure and model-serving architecture
  • High-growth AI company environment
  • Opportunity to work on AI systems being deployed by Fortune 500 organizations

Why Join

This is an opportunity to join a highly technical AI company working on one of the most important challenges in modern machine learning: making advanced language models significantly faster and more efficient.

You'll work directly on the infrastructure powering next-generation AI models, solving difficult problems across distributed systems, GPU computing, model inference, and production ML infrastructure.

As part of a small and technically elite team, you'll work closely with researchers and founders and have meaningful influence over how AI systems are designed, optimized, and deployed.

If you enjoy systems engineering, machine learning infrastructure, GPU optimization, high-performance inference, and solving difficult technical problems at the intersection of research and production, this role offers exceptional technical depth and impact.

Skills

PythonAWSAzureDockerKubernetesMachine LearningDeep LearningTensorFlowPyTorch

Similar Jobs

30

Senior Machine Learning Systems Engineer, Ads ML Experience Platform

Reddit · Remote - United States · Remote

3 days ago

Lead Business Analyst, Algo and Machine Learning Systems

Citicclsa · Hong Kong - One Pacific Place

1 week ago

Machine Learning Systems Engineer, Networking

Nvidia · US, CA, Santa Clara, United States of America

1 month ago

Machine Learning Systems Engineer, Ads ML Platform

reddit · Remote - The Netherlands +1 · Remote

1 month ago

Machine Learning Systems Engineer, Ads ML Platform

reddit · Remote - United Kingdom +1 · Remote

1 month ago

Senior Machine Learning Systems Engineer, Ranking Platform

reddit · Remote - United States · Remote

1 month ago

Machine Learning Systems Engineer, Networking

Nvidia · Santa Clara, CA,US, US

2 months ago

Python Software Engineer - Machine Learning Systems (m/f/d)

CUJU · Germany · Remote

3 months ago

Machine Learning Systems & Infrastructure Engineer

SpAItial · London +1 · Onsite

3 months ago

Scientist – Digital Discovery: Biological Data Systems & Machine Learning

Amgen is committed to unlocking · Canada - Burnaby · Onsite

3 months ago

Staff Machine Learning Systems Engineer

reddit · Remote - United States · Remote

4 months ago

Senior Machine Learning Systems Engineer

reddit · Remote - United States · Remote

4 months ago

Research Engineer, Machine Learning Systems

Deepgram · USA | Remote +2 · Remote

5 months ago

Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI

Scaleai · San Francisco, CA; New York, NY +1

9 months ago

Distinguished Engineer, Machine Learning Systems – Economy

Roblox · Remote

10 months ago

Machine Learning Systems Researcher

Lightmatter · Mountain View, CA

10 months ago

Principal Business Analyst, Algo and Machine Learning Systems, IT

Citicclsa · Hong Kong - One Pacific Place

11 months ago

Perennial Systems- Machine Learning Engineer

Nexthire · Remote · Remote

1+ year ago

Working Student - Machine Learning/Agentic Systems

ramblr.ai · Pullach im Isartal, Bayern

3 months ago

Systems Engineer - Machine Learning

General Robotics · Redmond, WA · Onsite

4 months ago

AI/Machine Learning Engineer (Embedded Systems, Inference Efficiency)

Qualcomm · Markham, ON,CA, CA

2 weeks ago

Senior Machine Learning Engineer, Agentic Systems - Moveworks

ServiceNow · Mountain View, CALIFORNIA, United States

3 weeks ago

Senior Machine Learning Engineer, Agentic Systems - Moveworks

ServiceNow · Mountain View, CALIFORNIA, United States

3 weeks ago

Machine Learning Engineer, Agentic Systems - Moveworks

ServiceNow · Mountain View, CALIFORNIA, United States

1 month ago

Senior Machine Learning Engineer (Large Systems)

Graphcore · Gdańsk, Pomeranian Voivodeship, Poland +3

1 month ago

Senior Machine Learning Engineer, Agentic Systems - Moveworks

ServiceNow · Mountain View, CALIFORNIA, United States

2 months ago

Senior Machine Learning Engineer, Agentic Systems - Moveworks

ServiceNow · Mountain View, CALIFORNIA, United States

3 months ago

Postdoctoral Fellow - Applied Machine Learning in Quantum Systems

Queracomputinginc · Boston, MA USA +1

5 months ago

Staff Machine Learning Engineer, Agentic Systems - Moveworks

ServiceNow · Mountain View, CALIFORNIA, United States

6 months ago

Staff Machine Learning Engineer - Recommendation Systems

Glance · Bangalore +1

7 months ago