Hiring.Camp

Site Reliability Engineer, Data Center Infrastructure

Spacex

·

Today

Location
Bastrop, TX · Bastrop, TX, United States
Type
Full-time
Department
Application Software
Experience
7+ years
Source
Greenhouse

Description

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

SR. SITE RELIABILITY ENGINEER, DATA CENTER INFRASTRUCTURE 

The application software team is the central nervous system of SpaceX. Manufacturing is how SpaceX turns designs into hardware. The compute, storage, and networking that run our factories must be as reliable as the products we build. This team owns infrastructure supporting Starship, Starlink, Starshield, and Terafab. This position will have a direct impact on factory uptime, throughput, and production scale across programs.

The ideal candidate has strong software engineering fundamentals and a passion for infrastructure: reliability, stability, proactive maintenance, and scalability. You understand the system before you change it, solve hard problems, communicate clearly with stakeholders and teammates, and take ownership of work that manufacturing depends on.

Aerospace experience is not required. We value smart, motivated, collaborative engineers who treat teammates with fairness, respect, and support, and who want to take full ownership of challenging problems to help make humanity multi-planetary.

RESPONSIBILITIES:

  • Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems across Starship, Starlink, Starshield, and Terafab
  • Manage infrastructure as code and use observability to provide a complete picture of platform health
  • Design for reliability, stability, and scale; find and remove bottlenecks with measurement and engineering
  • Practice proactive maintenance: capacity planning, lifecycle management, and reducing toil before it becomes an incident
  • Partner with software engineers, manufacturing stakeholders, and site teams to build operable, maintainable systems
  • Improve the full lifecycle—from design through deployment, operation, and continuous refinement
  • Practice sustainable incident response and blameless postmortems
  • Provide high-quality support to manufacturing and engineering users
  • Communicate clearly with stakeholders and teammates
  • Participate in on-call and travel to sites as needed for deployments, incidents, and cross-site reliability

BASIC QUALIFICATIONS:

  • Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 7+ years of professional experience in SRE or DevOps in lieu of a degree
  • 3+ years of experience with Python and Python-based development frameworks
  • Experience with Linux operating systems

PREFERRED SKILLS AND EXPERIENCE:

  • Experience with compute, storage, and/or networking infrastructure in production
  • Infrastructure as code (Terraform, Ansible, Puppet, or similar)
  • Containers and virtualization (Docker, Kubernetes, vSphere, QEMU, KVM, etc.)
  • Databases and data modeling (Postgres, Clickhouse, etc.)
  • Ability to translate high-level requirements into implementations from first principles
  • Comfort with mission-critical systems and appropriate urgency and care
  • Skillful communication with customers, peers, and management
  • Comfort operating across multiple sites and manufacturing programs

ADDITIONAL REQUIREMENTS:

  • Must be able to work extended hours and weekends as needed
  • Must be able to travel to different sites (Hawthorne, CA; Redmond, WA; Cape Canaveral, FL; Starbase, TX)
  • Ability to pass Air Force background check for Cape Canaveral 
  • This role requires you to be onsite. Remote and/or hybrid work will not be considered 

ITAR REQUIREMENTS:

  • To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.  

SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.

Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to [email protected]. 

Skills

PythonDockerKubernetesTerraformAnsibleLinuxDevOpsSRE

Similar Jobs

30

ServiceNow Platform Engineer

Lightfeatheriollc·Washington, DC +1·Hybrid

Today

DevOps Engineer

Lightfeatheriollc·US +1

Today

Platform Security Engineer

Scality·Paris, France·Onsite

Today

Site Reliability Engineer

Andurilindustries·Waltham, Massachusetts

Today

Senior Platform Engineer

CreateFuture·Sofia·Remote

Today

AV Engineer - Platform

Spotify·Stockholm·Onsite

Today

DevOps Engineer

Roboyo·London·Remote

Today

DevOps Engineer

Citi Bank·ONE PENNS WAY AMENITIES BLDG NEW CASTLE, US·Hybrid

1d ago

DevOps Engineer

Expleo Pt En·Lisbon, PT

1d ago

DevOps Engineer

Motorola Solutions·Cork, Ireland

1d ago

DevOps Platform Engineer

Booz Allen Hamilton·Rome, NY

1d ago

DevOps Engineer

citibank·New Castle, DE

1d ago

DevOps Engineer

Rbs·Edinburgh, UK +1·Hybrid

1d ago

Data Platform Engineer

Takeaway·Fleet Place Office, UK

1d ago

DevOps Engineer

motorolasolutions·Cork, Ireland

1d ago

Lead Platform Engineer

Mastercard·Pune, India·Hybrid

1d ago

DevOps Engineer

Infios·Bangalore, India

1d ago

Cloud Platform Engineer

Supabase·Remote, Global·Remote

1d ago

Data Platform Engineer

Accenture·Bengaluru, BDC7B

1d ago

DevOps Engineer

Accenture·Hyderabad, HDC3C

1d ago

Devops Engineer

citibank·Pune, MH

1d ago

Devops Engineer

Citi Bank·PLOT NO-1, S.NO. 77·Hybrid

1d ago

DevOps Engineer

Workwave·Atlanta, GA +1·Remote

1d ago

Staff Platform Engineer

Patientpoint·Remote, US·Remote

1d ago

DevOps Engineer

Alvaria·TX

1d ago

Site Reliability Engineer

Alvaria·TX

1d ago

Site Reliability Engineer

ACI Worldwide·Timisoara, Timis·Hybrid

1d ago

Site Reliability Engineer

ACI Worldwide·Timisoara, Timis·Hybrid

1d ago

DevOps Engineer

Movilges·Remote

1d ago

Senior Platform Engineer

GAINSCO·Richardson, TX

1d ago