Hiring.Camp

High Performance Compute (HPC) Software Engineer – HPC SW Systems

Kla

·

Yesterday

Salary
$106k – $155k
Location
USA-MI-Ann Arbor-KLA, United States of America
Type
Full-time
Department
Engineering
Education
PhD
Source
Workday

Description

To make electronics, you need chips, wafers, transistors, reticles, and... To make these, you must see, test and manufacture them at scale—faster and better than ever before. That's where KLA comes in. Whether you're early in your career or an experienced professional, you'll solve complex challenges, work alongside brilliant minds and help shape the future of technology.



Group/Division

The Semiconductor Products and Customers (Semi PC) business unit designs, builds and sells KLA’s product portfolio to help chip manufacturers meet high-quality standards for technologies like AI, data centers, automotive and electronic devices. Our inspection, metrology, and analytics systems detect and monitor defects across logic and memory chips, specialty semiconductors, wafers, reticles, advanced packaging, process equipment and materials. We also develop technologies for specialty processes, IC substrates, and PCB manufacturing, including etch and deposition, component inspection, and imaging and analytics. KLA's Central Engineering organization is made up of nine Centers of Excellence (CoEs) spanning disciplines like automation, motion control, sensors, platform design, and packaging. Each CoE contributes not only technical deliverables but also deep expertise, best practices, and advanced tools that elevate both what we build—and how we build it.

What You'll Do

In this role, you will play a key part in advancing business priorities by delivering high-impact work across your area of expertise.


Key Responsibilities

HPC Software Engineering

· Design, develop, and optimize HPC software running on large-scale Linux clusters, including distributed and parallel workloads (MPI, multithreading, GPU-accelerated pipelines, containerized workloads).

· Optimize application performance and power utilization across CPU, memory, storage, and network subsystem, with attention to throughput, latency, and scaling behavior.

· Develop and maintain system-level tooling for cluster bring-up, diagnostics, monitoring including component power usages, and health checks.

· Work closely with algorithms, systems and application teams to understand and translate workload characteristics into power-efficient HPC software solutions.

HPC Systems & Hardware Awareness

· Collaborate with hardware and systems teams to define HPC node, storage, and interconnect requirements based on software and algorithm needs.

· Understand and influence CPU/GPU selection, memory sizing, PCIe layout, NUMA behavior, and network topology to ensure optimal software performance.

· Participate in HW/SW co-debug activities, including performance bottlenecks, stability issues, and failure analysis.

Rack & Infrastructure Engineering

· Understand rack-level integration of HPC systems, focusing on power, cooling, cabling, networking, and physical layout considerations.

· Understand data-center and lab constraints such as power budgets, thermal limits, network drops, and serviceability.

· Contribute to best practices, and design reviews for new platforms and refresh cycles.

Cross-Functional Collaboration

· Act as a technical bridge between software, hardware, systems teams.

· Provide clear technical documentation covering software and system architecture, deployment flows, performance assumptions.

Required Qualifications

· Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.

· Strong experience developing HPC or systems software on Linux.

· Proficiency in Java and/or C++ and/or other system-level or performance-oriented languages.

· Hands-on experience with parallel computing (MPI, OpenMP, multithreading). Candidates with GPU computing (CUDA, ROCm, or equivalent) would be preferred.

· Solid understanding of HPC hardware fundamentals: CPUs, memory hierarchies, storage, networking (Ethernet / InfiniBand).

· Practical experience working with clusters, servers, or rack-scale systems in lab or production environments.

· Strong debugging skills across software, OS, and hardware boundaries.

Preferred Qualifications

· Experience with containerized HPC environments (Docker, Singularity/Apptainer, Kubernetes in HPC contexts).

· Familiarity with high-speed interconnects, storage architectures, and performance benchmarking.

· Exposure to rack integration, including cabling, power distribution, cooling, and system bring-up.

· Experience in semiconductor, manufacturing, or high-reliability systems environments.

· Ability to reason about system reliability, MTBF/MTBA, and failure modes in large compute installations.

What Makes This Role Unique at KLA

· Work on mission-critical HPC platforms that directly impact semiconductor manufacturing capability.

· Influence both software architecture and physical system design, not just code in isolation.

· Collaborate with world-class experts across algorithms, hardware, systems, and operations.

· See your work deployed at scale in real production tools—not just in the data center.

Minimum Qualifications

Doctorate (Academic) Degree and 0 years related work experience; Master's Level Degree and related work experience of 3 years; Bachelor's Level Degree and related work experience of 5 years

About KLA

We provide advanced inspection tools, metrology systems, process solutions, and computational analytics that make electronics possible, tackling complex challenges. From electron and photon optics to machine learning and data analytics, we seek perfection at the most fundamental level of matter in the universe. If you want to make electronics that push industries forward and make the world a better place, join us.



Total Rewards

Base Pay Range: $105,900.00 - $155,300.00 Annually Primary Location: USA-MI-Ann Arbor-KLA

KLA’s total rewards package for employees may also include participation in performance incentive programs and eligibility for additional benefits including but not limited to: medical, dental, vision, life, and other voluntary benefits, 401(K) including company matching, employee stock purchase program (ESPP), student debt assistance, tuition reimbursement program, development and career growth opportunities and programs, financial planning benefits, wellness benefits including an employee assistance program (EAP), paid time off and paid company holidays, and family care and bonding leave.

Interns are eligible for some of the benefits listed. Our pay ranges are determined by role, level, and location. The range displayed reflects the pay for this position in the primary location identified in this posting. Actual pay depends on several factors, including state minimum pay wage rates, location, job-related skills, experience, and relevant education level or training. We are committed to complying with all applicable federal and state minimum wage requirements where applicable. If applicable, your recruiter can share more about the specific pay range for your preferred location during the hiring process.



Use of AI Statement 

At KLA, our interviews seek to understand your individual skills, problem-solving approach and authentic thinking. To ensure a fair and consistent evaluation, the use of AI, recording tools or other technologies to generate, suggest or provide responses during interviews—whether virtual or in person—is not permitted unless explicitly approved in advance as part of a reasonable accommodation or invited by the interviewer. Use of these tools may interfere with our ability to evaluate your individual qualifications and affect your candidacy. KLA is committed to advancing innovation through responsible AI, and we value candidates who share this mindset.



Equal Opportunity Statement

KLA is proud to be an Equal Opportunity Employer. We will ensure that qualified individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us at [email protected] or at +1-408-352-2808 to request accommodation.


For additional information, view the US Know Your Rights poster on the U.S. Equal Employment Opportunity Commission website.

Skills

JavaDockerKubernetesLinuxMachine Learning

Similar Jobs

18

Senior Data Engineer - AI Infrastructure Integration, High Performance Compute

Ghr·New York, US +2·Hybrid, Onsite

1mo ago

Senior Data Engineer - AI Infrastructure Integration, High Performance Compute

Ghr·New York, US +2·Hybrid, Onsite

1mo ago

High Performance Compute Systems Site Lead (Onsite - LANL)

Hewlett Packard Enterprise (HP)·All, New Mexico·Remote, Hybrid, Onsite

2mo ago

High Performance Compute Systems Site Lead (Onsite - LANL)

Hewlett Packard Enterprise (HP)·All, New Mexico·Remote, Hybrid, Onsite

2mo ago

High Performance Compute Systems Site Lead (Onsite - LANL)

Hewlett Packard Enterprise (HP)·All, New Mexico·Remote, Hybrid, Onsite

2mo ago

High Performance Compute (HPC) Software Engineer – HPC SW Systems

Kla·USA-MI-Ann Arbor-KLA, US

3mo ago

High Performance Compute Responsible Engineer

Relativity·Long Beach, California +1

3mo ago

High Performance Compute (HPC) Software Engineer – HPC SW Systems

KLA·USA-MI-Ann Arbor-KLA, US

4mo ago

Principal Engineer Position for System Power for Advanced 2.5D/3D High-Performance Compute

Qualcomm·San Diego, CA

7mo ago

High Performance Compute Linux System Administrator

Hewlett Packard Enterprise (HP)·All, Mississippi·Onsite

1y+ ago

High Performance Compute Linux System Administrator

Hewlett Packard Enterprise (HP)·All, Mississippi·Onsite

1y+ ago

High Performance Compute Linux System Administrator

Hewlett Packard Enterprise (HP)·All, Mississippi·Onsite

1y+ ago

Director of Mass Market – High-Performance AI and Compute

Renesas Electronics·Austin, TEXAS

4d ago

Product Marketing Manager, High Performance AI and Compute Power

Renesas Electronics·Austin, TEXAS

1w ago

Director of Hyperscalers and Emerging AI – High-Performance AI and Compute

Renesas Electronics·Austin, TEXAS

2w ago

Associate Professor or DTU Tenure Track Assistant Professor in High-Performance Computing - DTU Compute

DTU·Kgs. Lyngby, DK

3w ago

Praktikum - System-Test bzw. System-Integration von High Performance Computern

Aumovio·Regensburg, BY

1w ago

[SDV2404]次世代統合ハイパフォーマンスコンピューター開発エンジニア/ Next Generation integrated High Performance Computer Development Engineer(一般層 総括職/担当職)

Alliance·Nissan Technical Center - Atsugi, Japan

7mo ago