Hiring.Camp

Inference Engineer

techire ai

·

May 21, 2026

Location
San Francisco, CA
Type
Full-time
Department
Engineering
Closing date
Today
Source
Vincere

Description

Machine Learning Engineer, Inference

Want to solve realtime inference problems where milliseconds genuinely matter?

This role is with a fast-growing voice AI company building the realtime speech infrastructure layer behind hundreds of millions of production conversations every month. Their systems power enterprise voice experiences used at massive scale across customer support, ordering, and conversational automation.

This is not another generic AI platform role focused on wrapping APIs or building dashboards.

The work here sits deep in the runtime stack, optimising realtime speech systems under production latency constraints. Think streaming inference, scheduler design, GPU utilisation, concurrency optimisation, dynamic batching, and making state-of-the-art speech models actually behave correctly in realtime environments.

You’ll join a lean engineering team working directly on the inference systems behind low-latency conversational speech models. The challenge is not simply generating outputs, it’s generating speech naturally, reliably, and fast enough for real human interaction.

Your work will include:

  • Building and optimising realtime TTS streaming infrastructure
  • Improving scheduler and batching systems for production workloads
  • Reducing TTFA/TTFB while maintaining speech quality and stability
  • GPU profiling and identifying kernel-level bottlenecks
  • Optimising TensorRT, Triton, ONNX Runtime, and custom serving systems
  • Managing KV cache systems, speculative decoding, and streaming inference
  • Supporting heterogeneous deployment environments across NVIDIA and AMD GPUs
  • Collaborating closely with model researchers to productionise cutting-edge speech systems

A large part of the role involves solving difficult runtime problems where latency consistency, concurrency, and throughput directly impact user experience. The team already operates beyond the performance of most publicly available realtime speech systems, but there’s still substantial room to push the infrastructure further.

You’ll likely have strong depth across inference systems, runtime optimisation, distributed serving, or GPU performance engineering. Experience with tools like TensorRT, Triton, vLLM, CUDA Graphs, ONNX Runtime, or custom schedulers would be highly valuable.

The environment suits engineers who naturally investigate bottlenecks, enjoy working close to hardware constraints, and care deeply about performance engineering. If reducing latency by 30ms feels meaningful, you’ll probably enjoy this team.

The stack includes Rust, C++, Python, CUDA, TensorRT, Triton, Kubernetes, AWS, and custom realtime inference infrastructure.

Compensation is highly competitive and flexible depending on experience, including strong salary, equity, and benefits.

Location: Remote across the US or Europe.

If you’re excited by realtime AI systems problems where optimisation work directly shapes production performance at scale, this would be worth exploring.

All applicants will receive a response.

Skills

PythonRustAWSKubernetesMachine Learning

Similar Jobs

30

Inference Engineer

Cartesia · *HQ - San Francisco, CA · Remote, Hybrid

1+ year ago

Sr. Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs

Amazon

Yesterday

Senior Software Engineer, Inference

AssemblyAI · Toronto · Remote

Yesterday

Senior Software Engineer, Inference

AssemblyAI · New York · Remote

Yesterday

Senior Software Engineer, Inference

AssemblyAI · Washington D.C. · Remote

Yesterday

Senior Software Engineer, Inference

AssemblyAI · Chicago · Remote

Yesterday

Senior Lead Software Engineer- Inference platform engineer

JPMorgan Chase · Seattle, WA, United States, US

2 days ago

Senior Lead Software Engineer- Inference platform engineer

JP Morgan Chase · Seattle, WA, United States, US

2 days ago

Research Engineer - Geo-Distributed Inference

Pluralis Research · USA or Australia · Remote

3 days ago

Research Engineer - Decentralized Training and Inference Verification

Pluralis Research · USA or Australia · Remote

3 days ago

Senior Software Engineer (AI Inference & Runtime Platform)

AZX · Seattle · Remote

4 days ago

Senior Software Engineer, Inference

Assemblyai · United Kingdom · Remote

5 days ago

Senior Software Engineer, Inference

Assemblyai · North America · Remote

5 days ago

Senior Inference Engineer, AGI

Amazon

5 days ago

Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon · Remote

5 days ago

Machine Learning Engineer, Causal Inference, Level 5

Snapchat · Santa Monica - 3340 Ocean Park Blvd, United States of America +4

6 days ago

Research Engineer - Inference

Eleven Labs · United Kingdom +3 · Remote

6 days ago

Software Engineer for Automotive AI Inference Stack (f/m/div)

Infineon · Stockholm, Stockholm County,SE, SE

6 days ago

AI Inference Engineer

Ffive · San Jose, United States of America +1

1 week ago

Senior Software Engineer - AI Inference Performance

Nvidia · US, CA, Santa Clara, United States of America

1 week ago

Senior Applied Scientist / Engineer, Training & Inference

Adobe · San Jose, United States of America

1 week ago

Senior Software Engineer - AI Inference Performance

Nvidia · Santa Clara, CA,US, US

1 week ago

Staff Applied AI Inference Engineer

Crusoe · Denver, CO - US · Onsite

1 week ago

Inference Performance Engineer

adaption · San Francisco +11 · Remote

1 week ago

Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together · London +1

1 week ago

Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together · India +1 · Remote

2 weeks ago

Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together · Amsterdam +1

2 weeks ago

Software Engineer, Inference (AI Data Engineering)

Spacex · Palo Alto, CA

2 weeks ago

DevOps Engineer (AI Inference)

Gcore · Cyprus, Germany · Remote

2 weeks ago

Staff Machine Learning Engineer, Causal Inference

Doordash USA · San Francisco, CA; Sunnyvale, CA; Los Angeles, CA; Seattle, WA; New York City, NY

2 weeks ago