Hiring.Camp

Staff Site Reliability Engineer (AI Platform)

Manychat

·

Today

Location
Barcelona, Spain · Amsterdam, North Holland, Netherlands
Workplace
Hybrid
Department
Engineering Org
Seniority
Senior
Experience
5+ years
Source
Greenhouse

Description

WHO WE ARE 🌍

We help creators get more out of every conversation with Instagram-focused automations and support for other channels like Messenger, WhatsApp, and TikTok. The result? Better engagement, more sales, and real, sustainable growth.

With a diverse team of 350+ people spread across three continents, we’re building the leading Chat Marketing platform that is used — and loved — by more than 1.5 million customers worldwide.


WHO WE'RE LOOKING FOR 🌟

We’re looking for a Senior Site Reliability Engineer who thrives at the crossroads of classic Linux and AWS infrastructure and modern Site Reliability Engineering. This is a high-impact, hybrid role designed for someone who can manage cloud resources, harden Kubernetes clusters, and shape a more reliable and developer-friendly platform.

We need you not just to maintain but to rethink and evolve our infrastructure, balancing hands-on operations with strategic improvements that future-proof our growing AI product landscape.

You’ll take over key responsibilities from our current Infra Lead who is transitioning to a software-focused role, giving you immediate ownership and space to shine.

WHY THE ROLE IS SPECIAL 💡

You won’t be a cog in a massive SRE org. You’ll be the bridge between Infrastructure and Engineering, shaping how we scale Kubernetes, how we approach platform reliability, and how developers ship fast without fear. You’ll get autonomy, ownership, and a smart, humble team excited to learn with you.

WHAT YOU’LL DO 🤖

  • Maintain and harden AWS infrastructure (EC2, ALB/NLB, WAF, IAM, CloudWatch)
  • Operate and evolve our EKS clusters powering Python-based AI services
  • Migrate existing services to Kubernetes using Terraform and Helm
  • Codify infrastructure with Terraform and manage host-level automation via Ansible
  • Build and improve CI/CD pipelines with GitHub Actions
  • Own observability efforts: Prometheus, Grafana, alerting, and on-call readiness
  • Support OS-level patching, certs, WAF rules, and general infra hygiene
  • Partner with engineers to guide best practices and drive platform reliability
  • Create clean, maintainable infrastructure documentation and playbooks
  • Occasionally support rare off-hours incidents (don’t worry, really rare)

TO SHINE IN THIS ROLE 💥

  • 5+ years of experience managing Linux in production (Ubuntu, Amazon Linux)
  • Strong experience with Kubernetes (ideally EKS), Helm, and Terraform
  • Comfort with running and debugging Python workloads in containers
  • Solid understanding of networking, IAM, and cloud security best practices
  • Hands-on Nginx experience (Ingress and reverse proxy setups)
  • Excellent communication skills; you can explain complex infra to devs clearly

NICE TO HAVE SKILLS 🛠️

  • Strong Ansible skills beyond the basics
  • PostgreSQL or Amazon RDS tuning and operations experience
  • Deep understanding of observability tools (Prometheus, Grafana, Loki, etc.)
  • Familiarity with PHP production environments
  • Experience with TDD, CI/CD best practices, and agile development
  • Any previous SRE-like exposure such as building resilience, automation, or incident tooling

WHAT WE OFFER 🤗
We care deeply about your growth, well-being, and comfort:

  • 🌍 Hybrid onboarding to start work remotely and relocation support for you and your family.
  • 💙 Comprehensive health insurance for both you and your family.
  • 📚 Professional development budget for conference tickets, online courses, and other relevant resources to help you grow.
  • 🫶 Flexible benefits package to tailor perks that matters most for you.
  • 🪴 Hybrid work and generous leave options to prioritize your work-life balance.
  • 🍽️ In-office perks, including free meals and snacks.
  • 🤝 Company-funded sport activities, annual offsites and team-building events.

 

Manychat is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance.

This commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your interview process to ensure you’re set up for success.

With my application, I accept the Manychat Privacy Policy.

 

Skills

PythonPHPAWSKubernetesTerraformAnsibleCI/CDLinuxNginxPostgreSQLGitHubSRE

Similar Jobs

30

Staff Site Reliability Engineer (AI Platform)

Manychat · Amsterdam, Netherlands +1

Today

Staff Site Reliability Engineer

Guidewire · Kuala Lumpur Office, Malaysia · Hybrid

1 week ago

Staff Site Reliability Engineer

Pingidentity · UK - Remote +1 · Remote

1 week ago

Staff Site Reliability Engineer

Ionq · Santa Clara, California, United States

1 week ago

Sr/Staff Site Reliability Engineer, Consumer Apps

Attain · Chicago, IL +1 · Remote, Hybrid, Onsite

2 weeks ago

Staff Site Reliability Engineer

Servicetitan · India Bengaluru, Karnataka

2 weeks ago

Staff Site Reliability Engineer

SimSpace Corporation · Remote - U.S. · Remote

2 weeks ago

Staff Site Reliability Engineer

Levi Strauss & Co. Careers · Bengaluru, India

2 weeks ago

Staff Site Reliability Engineer

Tenex · Remote, USA · Remote

2 weeks ago

Sr. Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

Zscaler · Hyderabad, IND

3 weeks ago

Staff Site Reliability Engineer

Stryker is one of the · Haryana, Gurugram International Techpark, Block I Phase 1 Floors G, 3, 4, 5, India · Hybrid

3 weeks ago

Staff Site Reliability Engineer

Chamberlain · Oak Brook, United States of America

3 weeks ago

Staff Site Reliability Engineer

Stackblitz · Remote · Remote

4 weeks ago

Staff Platform Engineer / Staff Site Reliability Engineer

Andurilindustries · Sydney, New South Wales, Australia

1 month ago

Staff Site Reliability Engineer - Paze

earlywarningservices · San Francisco, United States of America +2 · Hybrid

1 month ago

Staff Site Reliability Engineer

The Onset Jobs Marketplace · AU

1 month ago

Staff Site Reliability Engineer- Eng

UKG · Lowell, MA,US, US

1 month ago

Staff Site Reliability Engineer - Volcano

Kong · United States · Remote

1 month ago

Staff Site Reliability Operations Engineer

Calix will come from · Remote - USA, United States of America · Remote

1 month ago

Staff Site Reliability Engineer, Security

Stord · Remote, United States, United States of America · Remote

1 month ago

Sr Staff Site Reliability Engineer (SRE)

Arrow · Ahmedabad, India +4

1 month ago

Staff Site Reliability Engineer

Zoox · Foster City, CA · Hybrid

1 month ago

Staff Site Reliability Engineer (Collaboration Engineering)

NBCUniversal · Orlando, FL, United States · Hybrid

2 months ago

Senior Staff Site Reliability Engineer

Ironcladhq · San Francisco +2 · Hybrid

2 months ago

Sr Staff Site Reliability Engineer

Archer56 · San Jose, California, United States

2 months ago

Senior Staff Site Reliability Engineer

Hivewatch · El Segundo, CA +1

2 months ago

Staff Site Reliability Engineer - Site Experience

reddit · Remote - United Kingdom · Remote

2 months ago

Staff Site Reliability Engineer

Arcadia · Chennai, Tamil Nadu, India

2 months ago

Staff Site Reliability Engineer

Earnin · Mountain View, US +1 · Hybrid, Onsite

2 months ago

Staff Site Reliability Engineer

SimSpace Corporation · Remote - U.S. · Remote

2 months ago
Staff Site Reliability Engineer (AI Platform) at Manychat | Hiring.Camp