Hiring.Camp

HPC SYSTEMS ENGINEER

Uw

·

Today

Location
Seattle, Non-Campus, United States of America
Type
Full-time
Department
Engineering
Experience
4+ years
Education
Bachelor
Source
Workday

Description

Job Description

UW Information Technology has an outstanding opportunity for HPC Systems Engineer to join their team.

About this Opportunity
Reporting to Director, the HPC Systems Engineer is responsible for supporting the HPC efforts for the research computing group and related research computing technologies. This position also provids expertise and support to other endeavors when needed. This position requires a team-oriented professional, experienced in managing large IT systems for research using automated and software-defined approaches. This position regularly interfaces with other UWIT teams as well as research customers across campus. 
This position requires participation in a 24x7 on-call rotation.
This is an essential position is required to work remotely when the University suspends operations.

Key Duties:
25% Systems Engineering and Development: Design, develop, adapt, integrate, manage, optimize, or deploy monitoring for new or existing systems, software, or automation to continually improve the performance, reliability, recoverability, manageability, and usability of research cyberinfrastructure.

20% Systems Administration: Perform system management functions such as account administration, service administration, software installation and configuration, firmware updates, software updates, backing up and restoring data, improving system and service monitoring, or planning and execution of hardware replacements.

15% Operational Monitoring and Troubleshooting: Monitor HPC and related systems health and performance, and troubleshoot and resolve issues that impact performance, health, or reliability.

20% Technical Support: Monitor support ticket and alerting queue; triage, respond to, and resolve tickets and on-call pages as appropriate.

10% Systems Research, Evaluation and Architecting: Research and evaluate new and future technologies. Plan and architect new systems and the evolution and lifecycle of existing systems and clusters. Suggest or advise on new technologies to the Director of Research Computing Operations and other members of the Research Computing Operations team.
5% Training, Mentoring and Documentation: Effectively share experience and knowledge with other team members. 
5% Other duties as assigned Responsibilities

Required Qualifications
To be considered for this opportunity your application must demonstrate you meet both the minimum qualifications and additional qualifications listed below. Equivalent education and/or experience may substitute for minimum qualifications except when there are legal requirements, such as a license, certification, and/or registration.

Minimum Qualifications
•    Bachelors degree in computer science, information technology, scientific, engineering, or related field or experience
•    Minimum 4 years’ experience in Linux system administration experience or substantial experience working with Linux.
•    Basic knowledge of networking hardware, software, protocols, and concepts. Demonstrated ability to work with minimal supervision, both independently and as part of a team.
•    Demonstrated excellent written/oral communication skills, technical documentation skills, user liaison skills, and personal interaction abilities.
Applicants who do not meet these qualifications WILL NOT be forwarded to the Hiring Manager.

Preferred Qualifications
•    Knowledge of Containerization platforms (e.g., Docker, Apptainer, LXC, Podman).
•    Web development skills such as Python and Django.
•    Security compliance experience such as NIST 800-171, CMMC L2, and FAR to protect data such as HIPAA, CUI, Export Controlled, etc.
•    Progressively responsible experience as an engineer, architect, or role with comparable technical responsibilities in a large Linux HPC environment.
•    Extensive experience with administration of Linux operating systems in a production environment, including experience with Red Hat Enterprise Linux or derivatives such as CentOS or Rocky Linux.
•    Familiarity with SLURM or other HPC scheduler (PBS Pro, PBS/Torque, SGE/UGE, LSF, etc).
•    Experience designing, configuring, and troubleshooting networks using both Ethernet and high-performance interconnects such as Infiniband (e.g. experience with VLAN, MLAG, and LACP configurations).
•    Proficiency in programming/scripting languages in the context of systems engineering or administration, preferably including Bash, Golang, or Python.
•    Experience in the configuration and use of mass deployment tools such as Warewulf, MAAS, Foreman, xCAT, Cobbler, or similar.
•    Proficiency with the use of Git for source control in collaboration with a team with multiple contributors.
•    Ability to administer and troubleshoot large high-performance parallel filesystems such as IBM Storage Scale (GPFS), Lustre, BeeGFS, Ceph.
•    Experience with the use of configuration management tools such as Ansible, SaltStack, or Puppet.
•    Experience with monitoring, analysis and visualization of metrics, telemetry, and logs (e.g. Grafana, Prometheus, splunc elasticsearch, rsyslog, etc).
•    Experience in a data center environment (e.g., racking equipment, running cables, labeling, asset tracking). 
•    Experience with AI training and inference.
•    Experience with timeseries, relational, and non-relational database management or database query construction (e.g. influx, promql, redis, mongodb, postgres, mysql, nosql, mariadb, etc).
•    Scientific background, research experience, and/or experience in a University setting.

Working Conditions
Requires monitoring of e-mail and trouble ticket system for questions needing immediate response during business hours. 
On-call responsibilities for after-hours system outages. Server management will include both production (24x7x365) and development systems. 
Open office environment. Hybrid – expect to be in the office a minimum of two days per week.
This position requires participation in a 24x7 on-call rotation.
This is an essential position is required to work remotely when the University suspends operations.

About the Team
UWIT is the University of Washington’s central IT organization, supporting all three campuses, UW medical centers, and global research. We partner across the University to power teaching, learning, innovation, and discovery. We are here to serve UW. 
The Research Computing Operations group within the UWIT, designs, implements, and operates Linux-based High-Performance Computing (HPC) clusters and storage systems to offer UW researchers cost-effective yet powerful high-performance computing options that unlock new possibilities for their research.


 

Compensation, Benefits and Position Details

Pay Range Minimum:

$97,080.00 annual

Pay Range Maximum:

$157,764.00 annual

Other Compensation:

-

Benefits:

For information about benefits for this position, visit https://www.washington.edu/jobs/benefits-for-uw-staff/

Shift:

First Shift (United States of America)

Temporary or Regular?

This is a regular position

FTE (Full-Time Equivalent):

100.00%

Union/Bargaining Unit:

Not Applicable

About the UW 

Working at the University of Washington provides a unique opportunity to change lives – on our campuses, in our state and around the world.  

UW employees bring their boundless energy, creative problem-solving skills and dedication to building stronger minds and a healthier world. In return, they enjoy outstanding benefits, opportunities for professional growth and the chance to work in an environment known for its diversity, intellectual excitement, artistic pursuits and natural beauty. 

Our Commitment 

The University of Washington is committed to fostering an inclusive, respectful and welcoming community for all. As an equal opportunity employer, the University considers applicants for employment without regard to race, color, creed, religion, national origin, citizenship, sex, pregnancy, age, marital status, sexual orientation, gender identity or expression, genetic information, disability, or veteran status consistent with UW Executive Order No. 81.

To request disability accommodation in the application process, contact the Disability Services Office at 206-543-6450 or [email protected]

Applicants considered for this position will be required to disclose if they are the subject of any substantiated findings or current investigations related to sexual misconduct at their current employment and past employment. Disclosure is required under Washington state law

Skills

PythonDjangoDockerAnsibleLinuxMySQLMongoDBRedisElasticsearchGitComplianceHIPAAGo

Similar Jobs

29

HPC Systems Engineer

Northmark·Dallas - Victory Commons, US

5d ago

HPC Systems Engineer

KLA·USA-CA-Milpitas-KLA, US

3mo ago

HPC Systems Engineer

Radixexperienced·Chicago, Illinois +1

1y+ ago

Systems Engineer – HPC & GPU Infrastructure

" MAXISIQ, Inc."·Bethesda, MD·Onsite

3d ago

Mechanical Engineer – ADAS HPC (High-Performance Computing) Systems & Liquid Cooling (Orin/Thor) in Architecture and Network Solutions R&D | AUMOVIO Korea

Aumovio·Seongnam-si, Gyeonggi-do

1w ago

Senior Systems Engineer (HPC/Server Farm)

Cadence Design·SAN JOSE 07, US·Onsite

1w ago

Systems Engineer – HPC & GPU Infrastructure

Leidos·1662 Intelligence Community Campus - Bethesda MD, US·Onsite

2w ago

AI and HPC Systems Performance Engineer

Hewlett Packard Enterprise (HP)·Bengaluru, Karnātaka·Hybrid

1mo ago

AI and HPC Systems Performance Engineer

Hewlett Packard Enterprise (HP)·Bengaluru, Karnātaka·Hybrid

1mo ago

AI and HPC Systems Performance Engineer

Hewlett Packard Enterprise (HP)·Bengaluru, Karnātaka·Hybrid

1mo ago

Senior HPC and Quantum Systems Engineer

Nvidia·Westford, MA +3·Remote

1mo ago

Senior HPC and Quantum Systems Engineer

Nvidia·Westford, MA +3·Remote

1mo ago

Expert HPC Network/Systems Engineer (KEY-TE.HPC-03.080626)

Capitalsolutionsgroup·Ft. Meade, Maryland·Onsite

1mo ago

SR. SYSTEMS ENGINEER PRINCIPAL (HPC/AI System Administrator, Multi-discipline Expert)

GDIT·USA VA Home Office, US·Remote

1mo ago

Engineer II Systems (HPC)

Microchip·India - Chennai

2mo ago

Engineer II Systems (HPC)

Microchiphr·India - Chennai

2mo ago

Systems Engineer – HPC & GPU Infrastructure

Leidos·1662 Intelligence Community Campus - Bethesda MD, US·Onsite

2mo ago

Sr. High Performance Computing (HPC) Systems Engineer

Spacex·Starbase, TX +2

2mo ago

Sr. High Performance Computing (HPC) Systems Engineer

Spacex·Hawthorne, CA +2

2mo ago

High Performance Compute (HPC) Software Engineer – HPC SW Systems

Kla·USA-MI-Ann Arbor-KLA, US

3mo ago

Lead Systems Engineer (HPC)

Hub Princeton·Princeton, NJ

3mo ago

Lead Systems Engineer (HPC)

Main Princeton·Princeton, NJ

3mo ago

HPC/AI systems engineer

Hewlett Packard Enterprise (HP)·Bengaluru, Karnātaka·Hybrid, Onsite

4mo ago

HPC/AI systems engineer

Hewlett Packard Enterprise (HP)·Bengaluru, Karnātaka·Hybrid, Onsite

4mo ago

HPC/AI systems engineer

Hewlett Packard Enterprise (HP)·Bengaluru, Karnātaka·Hybrid, Onsite

4mo ago

Senior HPC Systems Engineer

GliaCell Technologies·Annapolis Junction, MD

4mo ago

High Performance Compute (HPC) Software Engineer – HPC SW Systems

KLA·USA-MI-Ann Arbor-KLA, US

4mo ago

Solutions Architect, HPC Systems Engineer

Nvidia·Remote

8mo ago

Senior Systems Administrator (HPC Lab Integration & Development)

Fuse Engineering Llc·Annapolis Junction, MD

1y+ ago