Hiring.Camp

LLM Inference & GPU Systems Consultant

Delan Associates, Inc

·

Yesterday

Location
Charlotte, NC, US
Workplace
Onsite
Type
Contract
Experience
8+ years
Source
Breezy HR

Description

Job Title: LLM Inference & GPU Systems Consultant

Location: Charlotte, NC (Onsite)

Duration: 6+ Months

Must be onsite at client in Charlotte, NC at least 3 days/week

Role Overview:

We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.

Key Responsibilities

NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.

Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.

Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.

Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.

Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.

Required Qualifications

8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.

8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).

Proficiency in OpenShift AI and GPU orchestration tools like RunAI.

Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.

Proven track record managing the Hugging Face deployment lifecycle.


Skills

Kubernetes

Similar Jobs

24

DL Performance Software Engineer - LLM Inference

Nvidia · Canada, Toronto

Yesterday

DL Performance Software Engineer - LLM Inference

Nvidia · Toronto, ON,CA, CA

Yesterday

Software Development Manager, LLM Inference Model Enablement, Neuron SDK

Amazon · Remote

1 week ago

Sr. Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone · New York, NY, United States of America +3

2 weeks ago

Senior Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone · New York, NY, United States of America +3

1 month ago

Principal LLM Inference Engineer

D Matrix · Santa Clara · Hybrid

1 month ago

AI Computing Software Development Engineer, LLM Inference

Nvidia · Shanghai, Shanghai,CN, CN +1

1 month ago

AI Computing Software Development Engineer, LLM Inference

Nvidia · China, Shanghai +1

1 month ago

Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone · New York, NY, United States of America +3

2 months ago

Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone · New York, NY, United States of America +3

2 months ago

Senior Deep Learning Researcher, LLM Inference

Nvidia · Tel Aviv-Yafo, Tel Aviv District,IL, IL

2 months ago

Senior Deep Learning Researcher, LLM Inference

Nvidia · Israel, Tel Aviv

2 months ago

Senior Deep Learning Research Engineer, LLM Inference

Nvidia · Israel, Tel Aviv

2 months ago

Senior Deep Learning Research Engineer, LLM Inference

Nvidia · Tel Aviv-Yafo, Tel Aviv District,IL, IL

2 months ago

Distributed LLM Inference Engineer

Anyscale · San Francisco +1 · Hybrid

3 months ago

Senior Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone · New York, NY, United States of America +3

4 months ago

LLM Inference Engineer (SR)

Job Board · Remote

4 months ago

Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone · New York, NY, United States of America +3

6 months ago

Senior Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone · McLean, VA, United States of America +2

6 months ago

Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone · New York, NY, United States of America +3

6 months ago

LLM Inference Frameworks and Optimization Engineer

Together · San Francisco, Singapore, Amsterdam +1 · Remote

1+ year ago

Machine Learning Scientist I/II, LLM Training & Inference Research

Lilasciences · Cambridge, MA USA; San Francisco, CA USA · Hybrid, Onsite

10 months ago

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

Plaud · San Francisco, CA · Hybrid

3 months ago

Distributed Training & Inference Optimization Engineer (LLM) - GPU Optimization Department (GPUOD)

Rakuten · Rakuten Crimson House, Japan · Remote, Hybrid

9 months ago