- Location
- Bangalore, India
- Type
- Full-time
- Department
- Engineering
- Seniority
- Lead
- Experience
- 8+ years
- Source
- Workday
Description
Overview
We are seeking a Lead Software Engineer with strong expertise in DevOps, Site Reliability Engineering (SRE), Cloud Platforms, Kubernetes, and GitOps practices. The ideal candidate will design, operate, and scale production Kubernetes platforms while driving reliability, automation, observability, and operational excellence across enterprise cloud-native applications.
This role will partner closely with Engineering, Product, Infrastructure, and Security teams to build resilient, Kubernetes-based cloud-native solutions and champion modern DevOps practices using ArgoCD, GitOps, and container orchestration best practices.
What You'll Do
DevOps & Platform Engineering
- Design, implement, and optimize CI/CD pipelines and Kubernetes-based deployment strategies.
- Build and maintain scalable, secure, and highly available cloud infrastructure.
- Drive infrastructure automation, configuration management, and platform standardization.
- Improve developer productivity through platform engineering and self-service capabilities.
Site Reliability Engineering (SRE)
- Establish and maintain reliability standards, SLAs, SLOs, and operational best practices.
- Lead incident management, root cause analysis, and reliability improvement initiatives.
- Enhance system performance, scalability, availability, and disaster recovery capabilities.
- Implement proactive monitoring, alerting, and observability solutions.
GitOps & Automation
- Implement and manage GitOps practices using ArgoCD for Kubernetes application delivery.
- Automate application deployments, environment provisioning, and configuration management.
- Ensure consistent, secure, and auditable deployments across all environments.
- Promote DevOps and GitOps best practices across engineering teams.
Kubernetes & Cloud Infrastructure
- Design, deploy, and manage production Kubernetes clusters across cloud environments (EKS, AKS, GKE, or equivalent).
- Build and optimize containerized application deployments using Docker and Kubernetes.
- Implement Kubernetes-native tooling including Helm, Kustomize, and operators for application lifecycle management.
- Configure and troubleshoot Kubernetes networking, ingress, storage, autoscaling, and resource management.
- Drive cluster reliability, upgrade strategies, capacity planning, and performance optimization.
- Partner with Security teams to implement Kubernetes security best practices (RBAC, network policies, pod security standards).
Collaboration & Leadership
- Partner with Engineering, Product, Infrastructure, and Security teams to deliver reliable solutions.
- Mentor engineers and provide technical leadership on DevOps and SRE best practices.
- Collaborate with global teams across EMEA and US regions.
- Be flexible to work across overlapping time zones and shifts when business needs require.
What We Are Looking For
- 8–10 years of experience in Software Engineering, DevOps, Platform Engineering, or Site Reliability Engineering.
- Strong hands-on experience with DevOps, SRE, CI/CD, and Infrastructure Automation.
- Experience implementing GitOps practices using ArgoCD.
- Strong expertise in Kubernetes cluster design, operations, and troubleshooting, including Docker, Microservices, and Cloud Platforms (AWS/Azure/GCP).
- Proven experience designing, deploying, and operating production Kubernetes clusters in enterprise cloud environments.
- Hands-on experience with Kubernetes tooling such as Helm, Kustomize, kubectl, and cluster management platforms.
- Experience with container orchestration patterns, microservices architecture, and cloud-native application design.
- Experience with Infrastructure as Code (Terraform or equivalent).
- Hands-on experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or similar platforms.
- Strong troubleshooting, production support, and incident management experience.
- Experience developing automation scripts using Python, Shell, or similar technologies.
- Excellent stakeholder management, communication, and leadership skills.
- Experience working with globally distributed teams across EMEA and US regions.
- Flexibility to support global stakeholders across different time zones.
Our Values
If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.