Hiring.Camp

SRE Engineering Manager - GPU Cloud

Scaleway

·

Jul 16, 2026

Location
Paris
Workplace
Hybrid
Type
Full-time
Department
GPU Cloud
Seniority
Manager
Source
Lever

Description

OUR STORY:
 
🇪🇺 Join Scaleway and shape the sovereign cloud of tomorrow !
Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.
 
Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.
 
With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.
 
Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.
 
📍 Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.
 

WHY WE NEED YOU ? 

Our growth is driving us to strengthen our GPU Cloud team to support our expanding infrastructure and key AI roadmap initiatives.

Your mission will be leading the Site Reliability Engineering (SRE) team in order to build, automate, and maintain a highly reliable, production-grade GPU cluster infrastructure powering our sovereign cloud.

YOUR FUTURE TEAM 

We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together.

You will be part of a team of 6 SREs within the GPU Cloud organization. The team focuses on critical AI and HPC infrastructure challenges, including automating key components of our stack and implementing support for modern GPU technologies.

YOUR DAILY ROUTINE 

Tasks

  • Lead and manage a team of 6 Site Reliability Engineers, supporting their career growth and technical execution

  • Design and implement automated solutions for server lifecycle management across GPU clusters

  • Design and implement observability, logging, and monitoring solutions for large-scale GPU clusters

  • Plan, prioritize, and manage the technical development roadmap for the SRE team

  • Collaborate and coordinate closely with software engineering, product, and cross-functional teams across Scaleway

  • Handle recruitment and career management for team members

  • Maintain, scale, and optimize high-availability production systems under heavy load

  • Participate in on-call rotations to ensure production reliability and fast incident resolution

ABOUT YOU 

HARDSKILLS:

  • Strong experience managing engineering teams in high-constraint production environments

  • Proven expertise with Kubernetes container orchestration

  • Direct experience with cluster management and virtualization tools (Proxmox, Warewulf)

  • Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana)

  • Exposure to modern GPU hardware ecosystems (Nvidia, AMD) and high-speed networking fabric (InfiniBand, Spectrum-X, Tomahawk)

  • Knowledge of distributed and high-performance storage solutions (Lustre DDN, VAST)

SOFT SKILLS:

  • Strong engineering leadership and team management capabilities

  • Technical rigor and high attention to detail in production-critical environments

  • Ability to handle high-pressure operational situations and manage incident stress pragmatically

  • Excellent communication skills with the ability to convey challenging messages effectively

  • Collaborative mindset with a focus on empowering engineers rather than micromanaging


WHAT YOU WILL FIND AT SCALEWAY ++++ 

  • Hybrid work: We offer up to 3 days of remote work per week.

  • Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.

  • Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.

  • Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.

  • International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.

  • Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.

🚀 Why join the Scaleway adventure?

✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.

✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.

✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.

🔜 THE NEXT STEPS …

  • Discovery call with HR
  • Technical interview with the HPC team to understand your technical skills and approach to the role
  • Manager interview to validate your expertise
  • Interview with an Engineering Manager / Head of Engineering to deepen discussions and assess your fit with the team 
  • HR interview and office visit to tour our offices and meet your future colleagues
 

Skills

KubernetesSRE

Similar Jobs

30

Manager, Site Reliability Engineering

Mastercard · O'Fallon, Missouri (Main Campus), United States of America

3 days ago

Manager, Site Reliability Engineering

Mastercard · Vancouver, Canada

4 days ago

Manager Site Reliability Engineering

Sabre · Poland - Poland -Tischnera · Hybrid

1 week ago

Engineering Manager, DevOps

Maven AGI · Boston · Hybrid

2 weeks ago

Engineering Manager, Site Reliability

Radar · New York +1 · Onsite, Remote

3 weeks ago

DevOps Engineering Manager

Fiserv is the global leader · Sunnyvale, California, United States of America · Onsite

3 weeks ago

Manager - Engineering/DevOps

Davies · Pune

3 weeks ago

Manager, Site Reliability Engineering

LSEG · Bucharest - Iuliu Maniu Boulevard, Romania

3 weeks ago

Engineering Manager (Site Reliability)

Moniepoint · Remote, India +1 · Remote

3 weeks ago

Engineering Manager (Site Reliability)

Moniepoint · Remote, Lagos, Nigeria +1 · Remote

3 weeks ago

Site Reliability Engineering Manager

NationsBenefits, LLC · Plantation, FL, US · Remote

4 weeks ago

Manager Site Reliability Engineering

Sabre · Poland - Poland -Tischnera · Hybrid

4 weeks ago

Manager, Site Reliability Engineering

Oracle · Reston, VA, United States, US

1 month ago

Manager, Site Reliability Engineering

Octanner · USA - Utah-Salt Lake City-Headquarters, United States of America

1 month ago

Site Reliability Engineering Manager

Cae · Krakow, Poland · Hybrid

1 month ago

Manager, Site Reliability Engineering

Docebo · Toronto, Ontario · Hybrid

1 month ago

Manager, Site Reliability Engineering

Mastercard · Pune, India

1 month ago

Manager, Site Reliability Engineering

PROS Holdings · BGR Sofia Hybrid, Bulgaria · Hybrid

1 month ago

Manager, Site Reliability Engineering

Mastercard · Toronto, Canada (Ethoca)

1 month ago

Engineering Manager, Site Reliability

Vanguard · USA - Majestic, United States of America

1 month ago

Engineering Manager, Site Reliability

Vanguard · USA - Majestic, United States of America

1 month ago

Engineering Manager, DevOps

Maintainx · San Francisco, Raleigh, Miami +4

3 months ago

DevOps Engineering Manager

CD Projekt RED · Warsaw, Masovian Voivodeship, Poland · Hybrid, Onsite

5 months ago

Manager, Site Reliability Engineering

Meredith · India-Bangalore-Ecoworld

6 months ago

Manager, Site Reliability Engineering

LSEG · IND-BLR-Divyasree Technopolis, India

8 months ago

Site Reliability Engineering Manager

Canonical · Home Based - APAC; Home based - EMEA · Remote

1+ year ago

Sr Manager, Site Reliability Engineering

Fis · IND PUNE FL7, India

Today

Senior Manager of Data Science Production Engineering, DevOps

Natera · San Carlos, CA

2 days ago

Senior Manager of Data Science Production Engineering, DevOps

Natera · US Remote · Remote

2 days ago

Senior Engineering Manager, DevOps & SRE – CIAM Platform

LSEG · IND-Bangalore-A, RMZ Infinity, India · Hybrid

3 days ago