Hiring.Camp

Software Engineer, Fleet Automation

Northmark

·

Today

Location
Dallas - Victory Commons, United States of America
Type
Full-time
Department
Engineering
Visa
Not sponsored
Source
Workday

Description

THE COMPANY

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients’ research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation.

THE POSITION

NMC² is seeking a Software Engineer to join the Fleet Automation team within the HPC & Infrastructure organization. This team owns the systems and tooling that keep hundreds of high-performance GPU compute nodes provisioned, configured, and operating at peak efficiency — spanning bare-metal provisioning, lifecycle management, and automated remediation at scale.

In this role, you will design and build the automation platforms, internal services, and APIs that allow NMC² to operate its growing fleet with speed and reliability. You will work at the intersection of software engineering and infrastructure — writing production-quality Go, C#, and TypeScript services that directly manage physical hardware, integrate with orchestration layers, and surface actionable observability to operations and on-call teams.

You will collaborate closely with Infrastructure Engineers, Network Engineers, and Research & Client teams to translate operational pain points into durable, maintainable automation. You will participate in on-call rotations and take ownership of system health through proactive monitoring, alerting, and incident response. The ideal candidate is equally comfortable architecting backend services and debugging Linux systems, thrives in ambiguous environments, and takes pride in eliminating manual toil through well-crafted tooling. This role is based in Dallas, TX out of our Victory Commons office.

RESPONSIBILITIES

  • Design, build, and maintain fleet automation services and internal platforms for provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute nodes.

  • Develop APIs and service integrations that enable Infrastructure and Operations teams to deploy, image, validate, and decommission hardware with minimal manual intervention.

  • Build and maintain backend services in Go, C#, and TypeScript with a strong focus on reliability, testability, and long-term maintainability.

  • Design and evolve data models and persistent state for automation workflows, working across relational and NoSQL databases as appropriate.

  • Build and maintain CI/CD pipelines that gate configuration changes, run automated hardware validation tests, and promote changes safely across environments.

  • Instrument systems for observability — designing metrics, alerts, and dashboards in Prometheus and Grafana that provide real-time fleet health visibility to on-call teams.

  • Participate in on-call rotations; own incident response, post-mortems, and follow-through on reliability improvements across the fleet.

  • Identify systemic gaps in fleet reliability and efficiency and champion engineering solutions that reduce operational toil at scale.

REQUIREMENTS

  • Bachelor’s Degree in Computer Science, Software Engineering, or equivalent practical experience.

  • 5+ years of software engineering experience building production backend services or infrastructure automation tooling.

  • Proficiency in Go, C#, or TypeScript.

  • Experience designing and working with relational and NoSQL databases to support stateful automation workflows and internal platform services.

  • Solid understanding of Linux systems — networking, storage, process management, and debugging on Ubuntu or RHEL variants.

  • Experience building and maintaining CI/CD pipelines and observability stacks (Prometheus, Grafana, Alertmanager, ELK) in a production environment.

  • Familiarity with GPU compute infrastructure and NVIDIA tooling (DCGM, nvidia-smi, NVIDIA Container Toolkit) is a strong plus.

  • Exposure to event-driven architectures or messaging platforms (e.g. Kafka) is a plus for teams building automation workflows across distributed services.

  • Strong communication skills and a collaborative mindset — comfortable navigating ambiguity, taking initiative, and working across Infrastructure, Operations, and Research teams.

It is impossible to list every requirement for, or responsibility of, any position.  Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company’s needs may change over time.  Therefore, the above job description is not comprehensive or exhaustive.  The Company reserves the right to adjust, add to or eliminate any aspect of the above description.  The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Benefits & Perks:

  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub

  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability

  • 401(k): Company will match 100% of your contributions up to 6%

  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.

  • Time Off:  25 days of Paid Time Off plus 12 company holidays


EQUAL OPPORTUNITY EMPLOYER

NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY'S POLICY IS NOT TO DISCRIMINATE AGAINST ANY APPLICANT OR EMPLOYEE BASED ON RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY OR EXPRESSION, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, AND GENETIC INFORMATION, OR ANY OTHER BASIS PROTECTED BY APPLICABLE LAW. THE FIRM ALSO PROHIBITS HARASSMENT OF APPLICANTS OR EMPLOYEES BASED ON ANY OF THESE PROTECTED CATEGORIES.

Skills

TypeScriptCI/CDLinux

Similar Jobs

27

Software Engineer - Fleet

Lambda · San Francisco Office (Fremont St) +2 · Hybrid

3 months ago

Senior Software Engineer, Fleet Telemetry & Control/Rust

Verse · San Francisco, CA +1 · Remote, Hybrid

2 days ago

Senior Software Engineer, Fleet Intelligence

Fleetio · Remote - USA, CAN, MEX · Remote

4 days ago

Software Engineer, Fleet Platform

Pathrobotics · Columbus, Ohio +1 · Hybrid

6 days ago

Software Engineer II - Fleet Enablement & Insights

Torcrobotics · Ann Arbor, MI, Blacksburg, VA, Fort Worth, TX +1

1 week ago

Lead Software Engineer (Frontend) Fleet

Corelight · North America

1 week ago

Software Engineer, Fleet Infrastructure

Andurilindustries · Boston, Massachusetts, United States; Washington, District of Columbia, United States +1

2 weeks ago

Software Development Engineer (AWS ML), Machine Learning Israel (MLIL) — FLOW sub-team (Fleet Lifecycle & Operational Workflows)

Amazon

2 weeks ago

Software Engineer, Fleet Simulation (Core Data Science)

Zoox · Foster City, CA · Onsite

2 weeks ago

Senior Systems Software Engineer - Fleet Debuggability

Nvidia · US, CA, Santa Clara, United States of America

3 weeks ago

Senior Systems Software Engineer - Fleet Debuggability

Nvidia · Santa Clara, CA,US, US

3 weeks ago

Senior Software Development Engineer (AWS ML), Machine Learning Israel (MLIL) — FLOW sub-team (Fleet Lifecycle & Operational Workflows)

Amazon

3 weeks ago

Senior Software Engineer, Fleet Maintenance

Fleetio · Remote - USA, CAN, MEX · Remote

4 weeks ago

Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

Nvidia · Santa Clara, CA,US, US

1 month ago

Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

Nvidia · US, CA, Santa Clara, United States of America

1 month ago

Software Engineer, Autonomous Fleet Orchestration

Andurilindustries · Costa Mesa, California, United States; Seattle, Washington, United States; Washington, District of Columbia, United States +1

2 months ago

Software Engineer - Ride and Fleet Services

Zoox · San Diego, CA +2 · Hybrid

3 months ago

Senior / Staff Software Engineer, Fleet Management

Viamrobotics · New York, NY +1 · Hybrid, Onsite

5 months ago

Senior/Staff Software Engineer - Ride and Fleet Services

Zoox · Foster City, CA · Hybrid

7 months ago

Senior Software Engineer, Fleet Orchestration and Optimization

Waymo · Hybrid

8 months ago

Software Engineer, Fleet Management

Openai · San Francisco · Hybrid

8 months ago

Software Engineer, Fleet Hardware Health

Openai · San Francisco · Remote

1+ year ago

Software Engineer (Fleet Management), London

Isomorphiclabs · London +1

1+ year ago

Senior Software Engineer, Server Fleet Infrastructure

Coreweave · Remote

1+ year ago

Software Engineer, Fleet Infrastructure

Openai · San Francisco · Hybrid

1+ year ago

Senior Java Software Engineer (m/w/d) für B2B SaaS Fleetwaro

Gps · Österreich · Onsite

6 months ago

Senior Java Software Engineer (m/w/d) für B2B SaaS Fleetwaro

Gps · Österreich · Onsite

6 months ago