Hiring.Camp

LLM Inference & GPU Systems Consultant

Delan Associates, Inc

·

Aug 6, 2026

Location
Charlotte, NC, US
Workplace
Onsite
Type
Contract
Experience
8+ years
Source
Breezy HR

Description

Job Title: LLM Inference & GPU Systems Consultant

Location: Charlotte, NC (Onsite)

Duration: 6+ Months

Must be onsite at client in Charlotte, NC at least 3 days/week

Role Overview:

We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.

Key Responsibilities

NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.

Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.

Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.

Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.

Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.

Required Qualifications

8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.

8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).

Proficiency in OpenShift AI and GPU orchestration tools like RunAI.

Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.

Proven track record managing the Hugging Face deployment lifecycle.


Skills

Kubernetes

Similar Jobs

20

AI Engineer 5 (FM Hosting, LLM Inference)

Capitalone·McLean, VA +3

1d ago

AI Computing Software Development Intern, LLM Inference - 2027

Nvidia·China, Shanghai +1

2d ago

AI Computing Software Development Intern, LLM Inference - 2027

Nvidia·Shanghai, CN +1

2d ago

Software Development Engineer, Alexa Excellence, Alexa LLM Inference, Capacity, & Efficiency

Amazon·Remote

2d ago

AI Engineer 5 (FM Hosting, LLM Inference)

Capitalone·San Jose, CA +3

1w ago

Software Engineer, LLM Inference

Nvidia·China, Beijing

1w ago

Software Engineer, LLM Inference

Nvidia·Shanghai, CN

2w ago

AI Systems Research and Development Engineer – LLM Inference Systems & Optimization

Snowflake·US-WA-Bellevue

2w ago

Software Engineer — Distributed LLM Inference Systems

Intel·CHN - Minhang, China

1mo ago

Software Engineer — Distributed LLM Inference Systems

Intel·CHN - Minhang, China

1mo ago

Software Development Manager, LLM Inference Model Enablement, Neuron SDK

Amazon·Remote

2mo ago

Principal LLM Inference Engineer

D Matrix·Santa Clara·Hybrid

2mo ago

Senior Deep Learning Researcher, LLM Inference

Nvidia·Tel Aviv-Yafo, IL

4mo ago

Senior Deep Learning Research Engineer, LLM Inference

Nvidia·Tel Aviv-Yafo, IL

4mo ago

Senior Deep Learning Researcher, LLM Inference

Nvidia·Israel, Tel Aviv

4mo ago

Distributed LLM Inference Engineer

Anyscale·San Francisco +1·Hybrid

4mo ago

LLM Inference Engineer (SR)

Job Board·Remote

5mo ago

Senior Lead AI Engineer (FM Hosting, LLM Inference)

Capitalone·McLean, VA +2

8mo ago

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

Plaud·San Francisco, CA·Hybrid

4mo ago

Distributed Training & Inference Optimization Engineer (LLM) - GPU Optimization Department (GPUOD)

Rakuten·Rakuten Crimson House, Japan·Remote, Hybrid

11mo ago