Hiring.Camp

HPC Infrastructure & Cluster Engineer

D2 Consulting

·

Today

Salary
$170 – $180
Location
Springfield, VA · East
Department
D2
Experience
5+ years
Clearance
Required
Source
Greenhouse

Description

**ACTIVE TS/SCI SECURITY CLEARANCE REQUIRED**

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment.  In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster.  Your primary mission is to ensure a highly available, secure, and optimized hardware foundation.  By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations

Basic Qualifications

  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand)
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g. Run:AI, SLURM)
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes
  • Experience writing automation and configuration scripts (e.g. Bash, Python) to streamline cluster maintenance
  • Proven ability to diagnose and resolve complex hardware, network, and OS-level issues

Preferred Qualifications

  • Familiarity with parallel file systems and high-throughput storage architecture
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies

Additional Information

  • All your information will be kept confidential according to EEO guidelines.
  • Compensation is unique to each candidate and relative to the skills and experience they bring to the position. The salary range for this position is typically $170-$180k. This does not guarantee a specific salary as compensation is based upon multiple factors such as education, experience, certifications, and other requirements, and may fall outside of the above-stated range.
  • Highlights of our benefits include Health/Dental/Vision, 401(k) match, Accrued PTO, STD/LTD/Life Insurance, Referral Bonuses, professional development reimbursement, and more!

D2 Technical Services is committed to a merit-based recruitment process and encourages applications from all qualified individuals.  As a Veteran-Owned Small Business, we particularly welcome applications from veterans who have the requisite skills and experience.  Job applicants that are interested in one of our openings and may require a reasonable accommodation to participate in the job application or interview process, should contact us to request an accommodation.

 

Skills

PythonKubernetesLinuxCompliance

Similar Jobs

22

Staff HPC Infrastructure Engineer

Gh · Remote-USA-CA, United States of America · Remote

6 days ago

AI HPC Infrastructure Engineer

Professional Analysisgroup · Boston, MA, US · Remote, Hybrid

1 month ago

EDA Computing / HPC Infrastructure MTS

Global Foundries · USA - Texas - Austin, United States of America +1 · Hybrid

1 month ago

IT Hardware Procurement Manager - AI/HPC Infrastructure

Armada · India (Remote) +1 · Remote

2 months ago

Category Manager- HPC Infrastructure

Northmark · Dallas - MacArthur Blvd, United States of America

5 months ago

Software Engineer, GPU Infrastructure - HPC

Openai · San Francisco · Remote

6 months ago

Staff Software Engineer, GPU Infrastructure (HPC)

Cohere · Canada +1 · Remote

7 months ago

Product Owner HPC Infrastructure

Nebul Bv

1+ year ago

Systems Engineer – HPC & GPU Infrastructure

Leidos · 1662 Intelligence Community Campus - Bethesda MD, United States of America · Onsite

1 month ago

HPC/ML Infrastructure Engineer

Spellbrush · San Francisco or Tokyo · Onsite

2 months ago

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together · San Francisco +1

2 months ago

Technical Specialist for Storage, Infrastructure and HPC

Cambridge Computer Services, Inc · Remote

7 months ago

Power Systems Architect, Expeditionary AI/HPC Data Center Infrastructure

Andurilindustries · Costa Mesa, California, United States

3 days ago

Linux System & HPC Administrator (IT Infrastructure & Operations)

Analogdevices · India, Bangalore, RMZ

3 weeks ago

HPC / AI Software Infrastructure Lead (E)

Kla · USA-MI-Ann Arbor-KLA, United States of America

2 months ago

HPC / AI Software Infrastructure Lead (E)

KLA · USA-MI-Ann Arbor-KLA, United States of America · onsite

2 months ago

Ingénieur Infrastructure IA Cloud & HPC (F/H)

JOBS · CRETEIL - CRE1, France

3 months ago

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

Nvidia · Poland, Warsaw +3 · Remote

1 month ago

Senior HPC Support Engineer - Ethernet and AI Infrastructure

Nvidia · TX,US, US · Remote

2 months ago

Senior HPC Support Engineer - Ethernet and AI Infrastructure

Nvidia · US, TX, Remote, United States of America · Remote

2 months ago

Senior HPC Support Engineer - Ethernet and AI Infrastructure

Nvidia · TN,US, US · Remote

2 months ago

Senior HPC Support Engineer - Ethernet and AI Infrastructure

Nvidia · US, TN, Remote, United States of America · Remote

2 months ago