Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
DiDi Research America
·Today
- Salary
- $170k – $351k
- Location
- San Jose, CA · Mountain View, California, United States
- Department
- Artificial Intelligence
- Seniority
- Senior
- Source
- Greenhouse
Description
About the Company
About The Role
We are seeking an experienced and mission-driven Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization to lead the performance tuning, deployment, and resource scheduling of cutting-edge AI models across on-vehicle and cloud infrastructure. In this role, you will design high-efficiency inference pipelines, build system-level stability frameworks, and optimize hardware execution to ensure ultra-low latency and rock-solid operational reliability. You will act as a technical leader in AI infrastructure, accelerating model iteration and bridging the gap between frontier deep learning algorithms and real-time autonomous systems.
Responsibilities
-
Own the deployment, optimization, and resource scheduling of vehicle-side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.
-
Lead vehicle-side system stability initiatives, conducting independent root-cause analysis and driving resolution for complex, system-level performance bottlenecks and runtime anomalies.
-
Architect and scale service-oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.
-
Track and evaluate cutting-edge industry methodologies, continuously integrating advanced optimization toolchains, quantization techniques, and execution engines.
-
Establish system-level profiling and telemetry frameworks using CUDA tools to monitor, analyze, and maximize hardware utilization across target GPU architectures.
-
Collaborate cross-functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.
Qualifications
-
Master’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related technical field.
-
3-8+ years of industry experience in high-performance computing, AI infrastructure, model optimization, or embedded deployment.
-
Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools.
-
Deep familiarity with mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM).
-
Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.
-
Demonstrated ability to diagnose complex software-hardware integration issues and drive scalable, production-grade solutions.
Preferred Qualifications
-
Hands-on experience optimizing and deploying AI models on the NVIDIA Thor platform, including hardware resource scheduling and acceleration.
-
Proven track record of serving large foundation models (e.g., LLaMA, Qwen, GPT) in production or high-throughput cloud pipelines using frameworks like vLLM, SGLang, TGI, or LightLLM.
-
Background in deep learning training frameworks (PyTorch) and practical experience with model quantization (INT8/FP8/AWQ), kernel fusion, or graph compilation.
-
Experience deploying real-time, high-availability AI workloads in autonomous vehicles, robotics, or edge devices.