Hiring.Camp

Engineer, Inference & Model serving

techire ai

·

Apr 30, 2026

Salary
$220k – $320k
Location
San Francisco, CA
Workplace
Remote, Onsite
Type
Full-time
Department
Engineering
Closing date
Today
Source
Vincere

Description

ML Model Serving Engineer

Want to build the layer that actually makes AI usable in real time?

You’ll join a team focused on inference, where performance is the product. This is about delivering low-latency, high-throughput systems across LLMs, speech, and vision models running in production, not offline experiments.

They’re building real-time AI systems that need to respond instantly, reliably, and at scale. That means solving hard problems around batching, GPU efficiency, memory constraints, and system-level bottlenecks that most teams never fully crack.

You’ll sit at the core of the platform, working across model serving, infrastructure, and performance optimisation. A big part of the role is pushing current tooling beyond its limits, extending frameworks, profiling bottlenecks, and designing systems that hold up under real-world load.

This is not about training models. It’s about making them fast, efficient, and production-ready.

What you’ll work on:

  • Building high-performance serving systems for LLM, speech, and vision models
  • Scaling inference to production workloads with strict latency requirements
  • Optimising GPU utilisation and execution efficiency
  • Implementing techniques like continuous batching, KV cache optimisation, speculative decoding, and prefill/decode separation
  • Improving frameworks such as vLLM, TensorRT-LLM, Triton, and SGLang
  • Profiling and debugging performance across GPU, memory, and system layers

What you’ll bring:

  • Strong experience with ML inference or model serving systems
  • Deep understanding of latency and throughput optimisation in production
  • Solid Python and PyTorch skills, plus a systems or performance engineering mindset
  • Familiarity with distributed systems and production infrastructure

Exposure to CUDA, GPU profiling tools, or systems like Kubernetes and Ray is useful, but the key is knowing how to make models run efficiently at scale.

You’ll join a highly technical team with experience across major AI labs and big tech. The environment is pragmatic, focused on solving real performance problems rather than abstract research.

There’s real ownership here. You’ll help define how next-generation AI systems are served.

Package:
$220,000 – $320,000 base + equity
San Francisco, onsite 3 days per week

If you’re interested in working on the part of AI that actually determines whether it works in the real world, this is worth exploring.

All applicants will receive a response.

Skills

PythonKubernetesPyTorch

Similar Jobs

30

Inference Engineer

Hyperbolic Labs·San Francisco, CA

2w ago

Inference Engineer

techire ai·San Francisco, CA

4mo ago

Inference Engineer

Cartesia·*HQ - San Francisco, CA·Remote, Hybrid

1y+ ago

AI Engineer 5 (FM Hosting, LLM Inference)

Capitalone·McLean, VA +3

2d ago

Software Engineer II - AI/ML, Neuron Inference

Amazon·Remote

3d ago

Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon·Remote

4d ago

Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon·Remote

4d ago

Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon·Remote

4d ago

Sr. Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon·Remote

4d ago

Software Development Engineer, Alexa Excellence, Alexa LLM Inference, Capacity, & Efficiency

Amazon·Remote

4d ago

Senior Staff Applied AI Inference Engineer

Crusoe·San Francisco, CA - US·Onsite

4d ago

Engineer, Staff-Embedded, AI inference ( preferred ) Release engineer

Qualcomm·Hyderabad, TS

4d ago

Intern - ML Inference Performance Engineer

Axelera·Eindhoven·Hybrid

5d ago

Operating Systems Engineer, On-Device Inference | Consumer Devices

Openai·San Francisco

5d ago

Senior System Software Engineer - Dynamo-Triton Inference Server

Nvidia·Santa Clara, CA +2·Remote

1w ago

Software Engineer - AI Inference for Science

Argonne·Lemont, IL USA·Remote, Hybrid, Onsite

1w ago

AI Engineer 5 (FM Hosting, LLM Inference)

Capitalone·San Jose, CA +3

1w ago

Senior System Software Engineer - Dynamo-Triton Inference Server

Nvidia·Santa Clara, CA +2·Remote

1w ago

Software Engineer, LLM Inference

Nvidia·China, Beijing

1w ago

Senior Software Engineer, Inference

Hewlett Packard Enterprise (HP)·Spring, Texas +2·Remote, Hybrid

1w ago

Principal Software Engineer, Inference

Hewlett Packard Enterprise (HP)·Spring, Texas +2·Remote, Hybrid

1w ago

Senior Software Engineer, Inference

Hewlett Packard Enterprise (HP)·Spring, Texas +2·Remote, Hybrid

1w ago

Principal Software Engineer, Inference

Hewlett Packard Enterprise (HP)·Spring, Texas +2·Remote, Hybrid

1w ago

Senior Software Engineer, Inference

Hewlett Packard Enterprise (HP)·Spring, Texas +2·Remote, Hybrid

1w ago

Principal Software Engineer, Inference

Hewlett Packard Enterprise (HP)·Spring, Texas +2·Remote, Hybrid

1w ago

AI Systems Engineer - Agents & Inference

Axelera·Netherlands +16·Hybrid, Remote

1w ago

Sr. Software Development Engineer, Inference Team - AWS Neuron

Amazon·Remote

1w ago

SDE, Neuron Inference, Neuron Inference

Amazon·Remote

1w ago

Senior Infrastructure Engineer (ML Inference)

Artificial.Agency·Edmonton +1·Remote

1w ago

Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

DiDi Research America·San Jose, CA +1

2w ago