Hiring.Camp

Senior/Staff Infrastructure Engineer

Fal

·

Feb 23, 2026

Salary
$180k – $250k
Location
San Francisco · San Francisco, California, United States
Department
Engineering
Seniority
Senior
Experience
5+ years
Source
Greenhouse

Description

You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including  provisioning, health monitoring, error detection, and recovery — and when something breaks that automation can’t fix, you drive resolution with partners.

Key responsibilities

  • Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc
  • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting
  • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals)
  • Leverage AI to an extreme level to build tools and automate alerting and recovery
  • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation
  • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage
  • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes)
  • Develop a suite of automated error detection and recovery processes
  • Work with partners to solve technical issues

Requirements

  • 5+ years experience managing bare-metal and VM server fleets at scale (100+ nodes)
  • Strong software engineering skills in Python; you write production tooling, not scripts
  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling
  • Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init
  • Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning
  • Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)
  • Experience building internal tools or dashboards for infrastructure visibility
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Nice to have

  • Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump)
  • Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2
  • Experience with AMD GPUs
  • Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM)
  • Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001)

Compensation

  • $180,000-250,000 plus equity + benefits

Location

  • San Francisco, CA

What we offer at fal

  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • We are currently hiring in downtown San Francisco.
  • We offer visa sponsorship and will help you relocate to San Francisco.
  • Health, dental, and vision insurance (US)
  • Regular team events and offsites

Skills

PythonTerraformAnsibleLinuxSOCComplianceProcurementSOC 2ISO 27001

Similar Jobs

30

Senior/Staff Infrastructure Engineer

Fal · Turkey · Remote

5 months ago

Senior / Staff Infrastructure Engineer

Knot · New York, NY

10 months ago

Senior / Staff Infrastructure Engineer

Apiphany · San Francisco · Hybrid

10 months ago

Sr. Staff Software Engineer, Systems Infrastructure

LinkedIn · Mountain View, CA, United States · Hybrid

Today

Senior Staff Software Engineer – Test Automation Infrastructure (ZIA Core)

Zscaler · San Jose, California, USA · Hybrid

Yesterday

Senior Staff, Software Engineer - Slack Front End Infrastructure

Slack · Georgia - Atlanta, United States of America +1 · Onsite

1 week ago

Senior Staff, Software Engineer - Slack Front End Infrastructure

Salesforce · Georgia - Atlanta, United States of America +1 · Onsite

1 week ago

Senior/Staff Software Engineer, Platform Infrastructure

Verkada · San Mateo, CA United States +1 · Onsite

1 week ago

Staff/Senior Software Engineer, Document Intelligence & Data Infrastructure

Helios Intelligence Platforms · New York City · Onsite

1 week ago

Senior Staff Software Engineer, Core Infrastructure

Robinhood · Bellevue, WA

1 week ago

Senior Staff Software Engineer, Core Infrastructure

Robinhood · Menlo Park, CA

1 week ago

Senior Staff Software Developer, Core Infrastructure

Robinhood · Toronto, Canada

1 week ago

Senior Staff Software Engineer, AI Infrastructure

LinkedIn · Mountain View, CA, United States · Hybrid

1 week ago

Principal / Staff / Senior Infrastructure Engineer

AllSpice · Boston +1 · Hybrid

1 week ago

Senior / Staff Software Engineer, Cloud & Real Time Infrastructure

Antora · Remote +1 · Remote

1 week ago

Sr. Staff Software Engineer, Systems Infrastructure

LinkedIn · Mountain View, CA, United States · Hybrid

2 weeks ago

Senior Staff Software Engineer, Systems Infrastructure

LinkedIn · Sunnyvale, CA, United States · Hybrid

2 weeks ago

Sr. Staff Software Engineer, AI Infrastructure

LinkedIn · Sunnyvale, CA, United States · Hybrid

2 weeks ago

Senior Staff Software Engineer, AI Infrastructure

LinkedIn · Sunnyvale, CA, United States · Hybrid

2 weeks ago

Sr. / Staff Software Engineer, Infrastructure (Autonomy)

DiDi Research America · San Jose, CA +1

2 weeks ago

Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering - Riyadh, KSA

Qualcomm · Riyadh, Riyadh Province,SA, SA

2 weeks ago

Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

DiDi Research America · San Jose, CA +1

2 weeks ago

Senior Staff Software Engineer, DC Infrastructure

Crusoe · San Francisco, CA - US +1 · Onsite

3 weeks ago

Senior Staff Engineer, CAE Infrastructure & R&D Software

Samsungsemiconductor · San Jose, California, United States · Onsite

3 weeks ago

Senior Staff Engineer - Infrastructure Engineering (Linux), OS Team

Geico · MD Bethesda Office, United States of America +3 · Hybrid

3 weeks ago

Senior/Staff Software Engineer, Search & Retrieval Infrastructure

Pinecone · New York City · Hybrid

4 weeks ago

Sr Staff Engineer - Verification, CI/CD Infrastructure & Embedded Hardware (AISW)

Qualcomm · Oakland, CA,US, US

4 weeks ago

CAD RTL/DV Infrastructure Engineer, Sr Staff

Qualcomm · Bengaluru, KA,IN, IN

1 month ago

Senior Staff/Staff Engineer - Fiat Payment Infrastructure

Okx · Hong Kong, Hong Kong SAR

1 month ago

Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms

Gitlab · Remote · Remote

1 month ago