Hiring.Camp

Senior/Staff Site Reliability Engineer

Fal

·

Feb 23, 2026

Salary
$180k – $250k
Location
San Francisco · San Francisco, California, United States
Department
Engineering
Seniority
Senior
Experience
5+ years
Source
Greenhouse

Description

You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing systems — from Kubernetes clusters to deployment pipelines to the networking layer that connects it all. You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better.

Key Responsibilities

  • Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads
  • Build and maintain CI/CD pipelines and deployment infrastructure
  • Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability
  • Build dashboards, alerting, and anomaly detection across our systems
  • Define and enforce SLOs and build out incident response processes
  • Manage and improve our networking, load balancing, and service mesh configurations
  • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering

Requirements

  • 5+ years experience in managing critical production systems and software development workflows
  • Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
  • Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
  • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
  • Proficiency in Python and either Go or Bash for tooling and automation
  • Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Nice to have

  • Experience with managing GPU and AI/ML workloads
  • Experience with kernel-based monitoring and routing (eBPF, XDP)
  • Experience with security tooling (Falco, Coroot, SIEM)
  • Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)
  • Experience with distributed storage systems (Ceph, Longhorn, etc.)

Compensation

  • $180,000-250,000 plus equity + benefits

Location

  • San Francisco, CA

What we offer at fal

  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • We are currently hiring in downtown San Francisco.
  • We offer visa sponsorship and will help you relocate to San Francisco.
  • Health, dental, and vision insurance (US)
  • Regular team events and offsites

Skills

PythonGoKubernetesTerraformAnsibleCI/CDLinuxSIEMSRE

Similar Jobs

30

Senior Staff Platform Engineer

Thehartford·Hartford CT- Home Office, US +3·Remote

1w ago

Senior Staff Platform Engineer

Equinix·Remote Location - PL, Poland +1·Remote

1mo ago

Senior Staff Platform Engineer

Equinix·POL Warsaw, Poland +1·Hybrid, Remote

1mo ago

Senior/Staff Platform Engineer

CodeRabbit·San Francisco·Hybrid

2mo ago

Senior / Staff Platform Engineer

Radar·New York +1·Onsite

2mo ago

Senior/Staff Platform Engineer

UpGuard·Melbourne +4·Remote

3mo ago

Senior/Staff Platform Engineer

VRChat·Anywhere·Remote

4mo ago

Senior / Staff Engineer - Platform

Sei Labs·New York City +1·Hybrid, Remote

5mo ago

Senior/Staff Software Engineer (Platform)

Renesas Electronics·Wroclaw, Poland·Hybrid

4d ago

Senior/Staff Software Engineer (Platform)

Renesas Electronics·Katowice, Poland·Hybrid

4d ago

Senior/Staff Software Engineer (Platform)

Renesas Electronics·Belgrade, Serbia·Hybrid

4d ago

Senior/Staff Software Engineer (Platform)

Renesas Electronics·Lisbon, Portugal·Hybrid

4d ago

Senior Staff DevOps Engineer

Gevernova·Noida, India·Hybrid

4d ago

Senior Staff Data Platform Engineer - Data Access Team

ServiceNow·Santa Clara, CALIFORNIA·Hybrid

5d ago

Senior/Staff Engineer, Liquidity Platform, Structured OTC

OKX·Hong Kong, Hong Kong SAR

5d ago

Senior/Staff Site Reliability Engineer

Factorial·A Coruña, ES·Onsite

5d ago

Staff / Senior Software Engineer, Security Fusion Platform

Anthropic·San Francisco, CA +1·Onsite

6d ago

Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors

Geico·MD Bethesda Office, US +3·Hybrid

6d ago

Senior/Staff DevOps Engineer

Twenty·New York, NY·Onsite

1w ago

Senior/Staff DevOps Engineer

Twenty·Arlington, VA·Onsite

1w ago

Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

ServiceNow·Kirkland, Washington·Hybrid

1w ago

Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

ServiceNow·Santa Clara, CALIFORNIA·Hybrid

1w ago

Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

ServiceNow·San Diego, CALIFORNIA·Hybrid

1w ago

Senior Staff AI Engineer - Agentic AI Platform (Remote Eligible)

Capitalone·San Francisco, CA +5·Remote

1w ago

Senior/Staff DevOps Engineer, Platform Infrastructure

Trase·Seattle, WA +1·Hybrid

1w ago

Senior/Staff DevOps Engineer, Platform Infrastructure

Redcellpartners·Seattle, WA +1·Hybrid

1w ago

Senior/Staff Site Reliability Engineer

Factorial·Barcelona, ES·Onsite

1w ago

Temporary DevOps/System Engineer (Sr.staff)

SBT Global·San Jose, CA

1w ago

Senior Staff Site Reliability Engineer

Nvidia·India, Bengaluru

1w ago

Senior/Staff Embedded Robotics AI Engineer – Embedded Platform

Renesas Electronics·Ho Chi Minh, Vietnam·Hybrid

1w ago