- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 7+ years
- Closing date
- Today
- Source
- Vincere
Description
Overview
Our client is a forward-thinking and innovative organization dedicated to building scalable, resilient, and secure infrastructure that empowers their developers and business teams. They are seeking a DevOps & SRE Engineering Lead to spearhead initiatives that ensure the reliability, scalability, and continuous delivery of our cloud services.
Responsibilities
Platform Reliability & Operations
- Own the availability, performance, and reliability of mission‑critical systems running on AWS and Alibaba Cloud
- Design and maintain high‑availability, multi‑region architectures with clear disaster recovery and failover strategies
- Define and enforce SLOs, SLAs, and error budgets, ensuring reliability is treated as a first‑class engineering objective
- Lead incident response, on‑call operations, root cause analysis, and post‑incident reviews
DevOps & Automation
- Drive infrastructure and deployment automation using Infrastructure as Code (IaC) and CI/CD pipelines
- Establish and continuously improve deployment strategies (blue‑green, canary, rollback)
- Reduce manual operations through tooling, scripts, and standardized workflows
- Improve deployment frequency while minimizing risk and downtime
Multi‑Cloud Strategy (AWS & Alibaba Cloud)
- Operate and optimize workloads across AWS and Alibaba Cloud, understanding and managing their respective strengths, constraints, and trade‑offs
- Define patterns and standards to ensure consistency, security, and reliability across cloud environments
- Partner with engineering teams to decide what runs where, balancing cost, performance, compliance, and resilience
- Manage cross‑cloud networking, DNS routing, and traffic management strategies
Observability & Monitoring
- Establish end‑to‑end observability across platforms, including metrics, logs, traces, and alerts
- Improve signal‑to‑noise ratio in alerting to reduce fatigue and accelerate incident response
- Implement proactive monitoring to identify risks before they impact users
Leadership & Collaboration
- Lead, mentor, and grow a team of DevOps / SRE / Platform engineers
- Work closely with application engineering, product, and business teams to align platform capabilities with delivery goals
- Communicate reliability risks and trade‑offs clearly to both technical and non‑technical stakeholders
- Champion a culture of ownership, continuous improvement, and blameless learning
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
- 7+ years of professional experience in DevOps, SRE, Platform Engineering or related roles.
-
Strong experience running AWS‑based production systems, plus hands‑on exposure to Alibaba Cloud
-
Proven history owning 24/7, high‑traffic or business‑critical platforms.
-
Deep knowledge of IaC (Terraform preferred), automation, and cloud‑native architecture.
- Solid understanding of Containerization using Docker and orchestration tools such as Kubernetes.
- Confidence in designing failover, DR, and traffic management strategies.
- A true SRE mindset — you care about outcomes, not just tools.
- The ability to explain risk and trade‑offs clearly to non‑technical stakeholders.
- Proven ability to mentor and lead teams, fostering a diverse and inclusive environment.
- Exceptional problem-solving skills and a proactive approach to identifying and addressing reliability challenges.
- Good command in both English and Chinese (Cantonese & Mandarin).