Hiring.Camp

GPU Performance Engineer | Experienced Hire

Sig

·

Mar 27, 2026

Salary
$200k – $300k
Location
New York, NY, US
Department
Engineering
Education
PhD
Closing date
3 weeks ago
Source
iCIMS

Description

Overview

We are looking for a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role is focused on workloads where off-the-shelf runtimes and vendor libraries do not fully exploit the structure of the model, and where custom kernels, memory layouts, and execution strategies can deliver meaningful gains.  

 

You will work closely with quantitative researchers and engineers to understand model structure, identify computational bottlenecks, and turn mathematical ideas into production-grade GPU implementations. You will use your understanding of GPU hardware to help shape models that are both mathematically effective and efficient to run. The problems span compact neural networks, tree-based models, and other structured inference workloads where latency, throughput, and efficiency all matter.  

 

This role is a strong fit for someone who enjoys low-level optimization, performance analysis, and translating abstract models into hardware-efficient code. 

 

What you'll do

  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads 
  • Develop fine-grained GPU implementations tailored to specific model structures
  • Analyze quantitative research models and computational bottlenecks to identify opportunities for parallelization and hardware-efficient execution 
  • Collaborate directly with quantitative researchers to translate mathematical models into high-performance compute pipelines 
  • Optimize end-to-end inference performance through kernel tuning, memory-layout design, execution strategy, I/O optimization, and precision tradeoffs 
  • Profile and benchmark GPU performance 
  • Improve latency and throughput in production inference systems 
  • Contribute to GPU architecture decisions and performance best practices 

What we’re looking for

  • Strong proficiency in writing and optimizing CUDA kernels 
  • Solid programming experience in C/C++ (preferred) 
  • Deep understanding of GPU architecture, including memory hierarchy, SIMT execution, occupancy, and latency/throughput tradeoffs
  • Ability to reason about numerical stability, precision, performance tradeoffs, and how model design choices affect hardware efficiency
  • Strong problem-solving skills and comfort working with low-level systems

 

Preferred qualifications 

  • PhD in mathematics, physics, computer science, engineering, or related quantitative field 
  • Strong background in linear algebra, probability, numerical methods, or scientific computing
  • Experience working with quantitative research teams or financial models
  • Demonstrated ability to improve real-world inference performance beyond baseline framework or library implementations
  • Familiarity with PTX-level behavior, tensor core utilization, or architecture-specific tuning 
  • Exposure to ONNX Runtime, TensorRT, Triton, TVM, or similar systems 
  • Exposure to neural networks, tree-based models (e.g., LightGBM), state space models (e.g., Mamba architectures), and experience with kernel fusion, custom operators, model compilation, or graph-level optimization

The annual base pay range for this role is $200,000 - $300,000 + discretionary bonus + benefits. Susquehanna considers factors such as scope and responsibilities of the position, work experience, education/training, key skills, as well as market and organizational considerations when extending an offer.

 

About Susquehanna

Susquehanna is a global quantitative trading firm powered by scientific rigor, curiosity, and innovation. Our culture is intellectually driven and highly collaborative, bringing together researchers, engineers, and traders to design and deploy impactful strategies in our systematic trading environment. To meet the unique challenges of global markets, Susquehanna applies machine learning and advanced quantitative research to vast datasets in order to uncover actionable insights and build effective strategies. By uniting deep market expertise with cutting-edge technology, we excel in solving complex problems and pushing boundaries together.

 

If you're a recruiting agency and want to partner with us, please reach out to [email protected]. Any resume or referral submitted in the absence of a signed agreement will not be eligible for an agency fee.

 

#LI-KH2

#LI-Onsite

Skills

Machine Learning

Similar Jobs

30

GPU Performance Engineer

Qualcomm · San Diego, CA,US, US

1 month ago

GPU Performance Engineer

Intel · USA - CA - Folsom, United States of America +1

4 months ago

GPU Performance Engineer

Intel · USA - CA - Folsom, United States of America +1

4 months ago

GPU Performance Engineer

Qualcomm · Hsinchu, Hsinchu City,TW, TW

9 months ago

Performance Engineer, GPU

Anthropic · San Francisco, CA | New York City, NY | Seattle, WA +1 · Onsite

10 months ago

GPU Performance Engineer

Genmo · San Francisco HQ

1+ year ago

Staff, GPU Performance Engineer

Samsung · 3655 N 1st St, San Jose, CA, USA, United States of America +1 · Onsite

1 month ago

Senior System Software Engineer - GPU Performance Profiling Tools

Nvidia · China, Shanghai

1 month ago

Senior System Software Engineer - GPU Performance Profiling Tools

Nvidia · Shanghai, Shanghai,CN, CN

1 month ago

GPU Performance Engineer - Neural Reconstruction

Nvidia · Canada, Remote · Remote

1 month ago

GPU Performance Engineer - Neural Reconstruction

Nvidia · CA · Remote

1 month ago

Software Engineer, GPU Performance Tools

Nvidia · US, CA, Santa Clara, United States of America +2 · Remote

1 month ago

Software Engineer, GPU Performance Tools

Nvidia · Santa Clara, CA,US, US +2 · Remote

1 month ago

GPU Performance Engineer - Neural Reconstruction

Nvidia · US, CA, Remote, United States of America +5 · Remote

1 month ago

GPU Performance Engineer - Neural Reconstruction

Nvidia · CA,US, US +5 · Remote

1 month ago

Senior Systems Software Engineer - GPU Performance at Scale

Nvidia · Switzerland, Remote +1 · Remote

2 months ago

Senior Systems Software Engineer - GPU Performance at Scale

Nvidia · CH +1 · Remote

2 months ago

AI Frameworks Engineer – GPU Performance for Generative AI (OpenVINO)

Intel · KOR - Seoul, Korea, Republic of

2 months ago

AI Frameworks Engineer – GPU Performance for Generative AI (OpenVINO)

Intel · KOR - Seoul, Korea, Republic of

2 months ago

PRINCIPAL ENGINEER, GPU PERFORMANCE, SMAI

Micron · Taichung - AATT, Taiwan +1

4 months ago

PRINCIPAL ENGINEER, GPU PERFORMANCE, SMAI

Micron Technology · Taichung City,TW, TW +1

4 months ago

Senior System Software Engineer - GPU Performance

Nvidia · US, CA, Santa Clara, United States of America +1 · Remote

4 months ago

Senior System Software Engineer - GPU Performance

Nvidia

5 months ago

Senior Engineer, GPU Performance Architect (PPA)

Samsung · 3655 N 1st St, San Jose, CA, USA, United States of America +1 · Onsite

5 months ago

Data Center GPU Performance Engineer – Product

Nvidia

6 months ago

Member of Technical Staff (GPU Performance Engineer)

Reka · US, UK, Singapore, Remote · Remote

6 months ago

Member of Technical Staff - GPU Performance Engineer

Liquid Ai · San Francisco +2 · Remote

11 months ago

GPU Performance Verification Engineer

Qualcomm · Santa Clara, CA,US, US

2 weeks ago

GPU Performance Verification Engineer - Cork, Ireland

Qualcomm · Cork, CO,IE, IE

1 month ago

GPU Performance Verification Engineer

Intel · SRR4 - SRR4 - Sarjapur 4, India

1 month ago