Hiring.Camp

Site Reliability Engineer (Amsterdam)

Together AI

·

Apr 28, 2025

Location
Amsterdam · Amsterdam, North Holland, Netherlands
Department
Engineering
Experience
7+ years
Education
Bachelor
Source
Greenhouse

Description

As a Site Reliability Engineer (SRE) at Together, you are responsible for keeping all user-facing services and production systems running smoothly. You are a blend of a pragmatic operator and a software engineer that applies sound engineering principles, operational discipline, and mature automation to our operating environments and codebase.

You specialize in systems (operating systems, storage subsystems, networking), while implementing best practices for availability, reliability and scalability, with varied interests in algorithms and distributed systems.

Requirements

  • 7+ years of professional SRE or related experience
  • Bachelor's degree in Computer Science or a related field or equivalent work experience
  • Expert knowledge of Ansible (roles, playbooks), Terraform, and Kubernetes
  • Proficiency in programming/scripting languages
  • Direct experience in monitoring and observability practices
  • Advanced knowledge of cloud services
  • Ability to thrive in a collaborative environment involving different stakeholders and subject matter experts

Responsibilities

  • Be on an on-call (PagerDuty) rotation to respond to incidents that impact availability
  • Build and run our infrastructure with Ansible, Terraform, and Kubernetes to enable scaling to a massive number of concurrent users
  • Build monitoring systems to ensure the highest quality service for our customers
  • Design and implement operational processes (such as deployments and upgrades)
  • Debug production issues across all services and levels of the stack
  • Identify improvements for the product architecture from the reliability, performance and availability perspectives 
  • Plan the growth of Together AI’s infrastructure

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  

Skills

KubernetesTerraformAnsibleSRE

Similar Jobs

30

DevOps Engineer

BeVera Solutions LLC · Atlanta, GA

Today

Senior Platform Engineer

Natera · US Remote · Remote

Today

Senior Platform Engineer

Focused · Chicago, Illinois, United States +1

Today

Site Reliability Engineer

Bay Systems Consulting · Berkeley, CA

Today

Site Reliability Engineer

OnBoard · United States

Today

DevOps Engineer

Sprezzatura Management Consulting · Remote, US · Remote

Today

Devops Engineer

Miratech · Bengaluru, KA, India · Remote

Today

DevOps Engineer

Inetum · Lisbon, Lisbon, Portugal · Hybrid

Today

Release Platform Engineer

bet365 · Manchester, England, United Kingdom · Hybrid

Today

Release Platform Engineer

bet365 · Stoke-on-Trent, England, United Kingdom · Hybrid

Today

Site Reliability Engineer

Fortive · IN +1 · Remote

Yesterday

Site Reliability Engineer

Fluke · Remote, India · Remote

Yesterday

DevOps Engineer

Capital Markets Gateway · Brno · Hybrid

Yesterday

DevOps Engineer

Agdata · Pune, Maharashtra

Yesterday

Site Reliability Engineer

Andurilindustries · Waltham, Massachusetts, United States

Yesterday

Site Reliability Engineer

TJX · CAN Home Office Mississauga ON, Canada

Yesterday

DevOps Engineer

Swbc · SWBC Headquarters, United States of America

Yesterday

Engineer - DevOps

Allegion · Bangalore, India

Yesterday

Site Reliability Engineer

Truist · Atlanta GA - 303 Peachtree Center Avenue - Garden Offices, United States of America +2 · Remote, Hybrid

Yesterday

Site Reliability Engineer

Truist · Atlanta GA - 303 Peachtree Center Avenue - Garden Offices, United States of America +2 · Remote, Hybrid

Yesterday

Power Platform Engineer

RSM · SLV-San Salvador-Calle Cortez Blanco #8 Urb. Madreselva, El Salvador

Yesterday

Devops Engineer

Playpower Labs · Fully Remote - India Only · Remote

Yesterday

Staff Platform Engineer

Ftmo · Prague office · Onsite

Yesterday

DevOps Engineer

G2I · USA

Yesterday

Technology Platform Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Technology Platform Engineer

Accenture · Chennai, CDC2F, India

Yesterday

Senior Platform Engineer

Amadeus · Lisbon - Carnaxide, Portugal · Hybrid

Yesterday

Cloud Platform Engineer

Accenture · Pune, PDC5A, India

Yesterday

Cloud Platform Engineer

Accenture · Bhubaneswar, BBDC1A, India

Yesterday

Data Platform Engineer

Accenture · Bengaluru, BDC9A, India

Yesterday
Site Reliability Engineer (Amsterdam) at Together AI | Hiring.Camp